Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D
MTP-3 restored prompt caching. Our deeper speculative mode cleared the cache for correctness. Using ordinary state backups at depth 3 let follow-ups reuse it. Better CPU kernels unpack IQ4_XS weights once for several speculative tokens, reducing repeated work. Lower-memory GPU attention processes a…
Read the full story at r/LocalLLM ↗
Timeline · 4 reports
- 2026-09-20 03:53 · r/LocalLLM
Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3) - 2026-09-20 03:44 · r/LocalLLaMA
Finally got Qwen 3.8 Next running on my v100 6gpu setup (TP2 PP3) - 2026-09-19 19:44 · r/LocalLLaMA
Finally got qwen 3.8 q4 k_M 27b runnin slick 40tkps load at 240k inf 4qnl for stable agentic coding on a 7900xtx. I have been finally able to give hermes a task and come back to results. Built a panel to manage my inference servers from hermes. - 2026-09-18 22:38 · r/LocalLLM
Qwen q4 3.8 27b 16 tok/s 32k RTX 3060 :D