AINewsnow

On a Tesla P40, loading a diffusion model in fp16 made it 14x slower than fp32

I run the AI for a small side project, Fring , a free wardrobe app, on a second-hand Tesla P40 in my homelab. The virtual try-on uses Leffa, a diffusion model. For a month every render took about 21 minutes, and I assumed that was just what a 2016 card could do. It wasn't. The model was loaded in f…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-10-09 08:11 · DEV Community — Machine Learning
    On a Tesla P40, loading a diffusion model in fp16 made it 14x slower than fp32

More stories

  1. Tesla Powerwall and Car Backup During an Outage — Matthew Berman
  2. Is the Tesla V100 still a valid card for a local inference host? — r/LocalLLM
  3. Are Tesla K80s any good for inference? — r/LocalLLaMA
  4. Is it still worth buying a Tesla V100 for machine learning in 2026? — r/learnmachinelearning
  5. GPT-6 and Intelligent UI for everyone — OpenAI News
  6. Introducing Mistral Large 4 — Mistral AI News
  7. Introducing Claude Haiku 5.5 on AWS — AWS Machine Learning Blog
  8. Sharing AI progress in mathematics — OpenAI News

Get the daily brief of stories like this at 6:30 every morning →