~75 tok/s in rtx 3090 with 90k Context, Q4-UD_K_XL [Qwen 3.8 27B]
This story is from 2026-08-23. It is preserved in the archive; the latest stories are on the live feed.
Been messing around with Qwen 27B ( Qwen3.8-27B-UD-Q4_K_XL.gguf ) and the separate MTP draft module in llama.cpp over the past few days. At first my speeds were either barely matching baseline (~50 t/s) or dropping down to ~35 t/s, but after tweaking flags and isolating bottlenecks, I finally got i…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-23 04:30 · r/LocalLLM
~75 tok/s in rtx 3090 with 90k Context, Q4-UD_K_XL [Qwen 3.8 27B]