48 tg/s 440 prefill on my grandma's cluster (2xP40) (sort of)
This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.
TL;DR: switching KV cache to f16 may give a boost in speed if using MTP and ngrams. I have a self-built "AI mega-cluster" with 2x P40s on a cheap Chinese motherboard and a Xeon CPU (around $1,100 to build, including water cooling for the GPUs). I was normally getting up to 15 tk/s with Qwen 3.8 27B…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-06 10:38 · r/LocalLLaMA
48 tg/s 440 prefill on my grandma's cluster (2xP40) (sort of)