On a Tesla P40, loading a diffusion model in fp16 made it 14x slower than fp32
I run the AI for a small side project, Fring , a free wardrobe app, on a second-hand Tesla P40 in my homelab. The virtual try-on uses Leffa, a diffusion model. For a month every render took about 21 minutes, and I assumed that was just what a 2016 card could do. It wasn't. The model was loaded in f…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-10-09 08:11 · DEV Community — Machine Learning
On a Tesla P40, loading a diffusion model in fp16 made it 14x slower than fp32