AINewsnow

Older CUDA Cards (p40 24gb) 25-30 t/s Qwen3.8-27B-MTP-BF16.gguf 54gb quantized to 18.4gb

This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.

Was looking into building a quantization method to leverage the add+multipy one clock cycle on my older Tesla p40 ($250 card, HP Z4G4 64gb xen-2133 $400). During the process I learned more about quantization methods and performance hits the older cards take with the newer models. In a nutshell, any…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-05 07:17 · r/LocalLLM
    Older CUDA Cards (p40 24gb) 25-30 t/s Qwen3.8-27B-MTP-BF16.gguf 54gb quantized to 18.4gb

More stories

  1. Migrating from a R640 with passive cards — r/LocalLLaMA
  2. V100 16GB worth it?? — r/LocalLLM
  3. Sunsetting the NVIDIA Tesla P100 GPU on September 15, 2026 | What will happen to these P100, can we buy them? — r/LocalLLM
  4. When do you think we’ll get physical AGI? — r/singularity
  5. Seasonic Vertex GX-1200W exploded.... — r/LocalLLM
  6. Anthropic says Claude 'leads' 26 percent of its AI R&D work — Engadget
  7. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  8. Introducing Kimi K3 on Amazon Bedrock — AWS Machine Learning Blog

Get the daily brief of stories like this at 6:30 every morning →