GPU drops to idle clocks during token generation with MTP / speculative decoding, help?
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
I can't figure it out. I'm running Unsloth's Qwen 3.8 27B Q4_K_S with their Q4 MTP model on a local Llama-server with a 3090. During prefill, the gpu gets the full memory and clock speeds, but when it comes time for token generation, the memory and clock speeds drops by half. The only thing that wo…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-11 10:18 · r/LocalLLM
GPU drops to idle clocks during token generation with MTP / speculative decoding, help?