I pushed Qwen3.8-27B to 381 tps for a single request on a RTX 3090
This story is from 2026-08-20. It is preserved in the archive; the latest stories are on the live feed.
Four days ago I released a hyper-optimized Qwen3.8-27B inference engine for an RTX 3090 (82 tps single request, 672 peak). Since then it went to ~114, then ~138 tps single-user with DFlash2 drafting and lookup-augmented drafting. Today it's ~133 tps on real chat prompts, 382 tps when the model repr…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-20 20:15 · r/LocalLLaMA
I pushed Qwen3.8-27B to 381 tps for a single request on a RTX 3090