Qwen3.8-27B at 147 tok/s on one CMP 170HX — 65 tok/s at 250K CTX
I’ve been tuning Qwen3.8-27B + DFlash2 on a single NVIDIA CMP 170HX 64GB (SM80) with vLLM 0.27.1. I’m finally calling the main optimization path done. Results Single-stream decode, greedy, fresh engine per cell, 1350 MHz / 180W : Context Peak decode 4K 134.7 tok/s 32K 147.0 tok/s 65K 100.1 tok/s 12…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-20 07:06 · r/LocalLLM
Qwen3.8-27B at 147 tok/s on one CMP 170HX — 65 tok/s at 250K CTX