Older CUDA Cards (p40 24gb) 25-30 t/s Qwen3.8-27B-MTP-BF16.gguf 54gb quantized to 18.4gb
This story is from 2026-09-05. It is preserved in the archive; the latest stories are on the live feed.
Was looking into building a quantization method to leverage the add+multipy one clock cycle on my older Tesla p40 ($250 card, HP Z4G4 64gb xen-2133 $400). During the process I learned more about quantization methods and performance hits the older cards take with the newer models. In a nutshell, any…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-05 07:17 · r/LocalLLM
Older CUDA Cards (p40 24gb) 25-30 t/s Qwen3.8-27B-MTP-BF16.gguf 54gb quantized to 18.4gb