Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070
This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.
ik_llama.cpp merged qwen4exp MTP support yesterday (PR #2369, mine, reviewed and tested by four other people on their own hardware). It's on main now, no fork or patch needed. Posting since the last couple threads had people saying MTP for this model only exists as an unsloth fork PR... there's ano…
Read the full story at r/LocalLLaMA ↗
Timeline · 5 reports
- 2026-09-05 19:53 · r/LocalLLM
Qwen3.8-Flash-Next workcase part 2: after the self-portrait, a painting and a map sheet, both unattended on a 3060 - 2026-09-05 12:52 · r/LocalLLaMA
Qwen3.8-Flash-Next (UD-Q4_K_XL) on a single RTX 3090 24GB + 128GB DDR4, is this config optimal? - 2026-09-05 08:08 · r/LocalLLM
Qwen3.8-Flash-Next drew a self-portrait site from one prompt on an RTX 3060, then signed it “Claude” - 2026-09-04 15:21 · r/LocalLLM
Latest llama.cpp vs experimental MTP build: Qwen3.8 Flash Next coding task on M5 Max — 9m24s vs 5m18s - 2026-09-03 16:30 · r/LocalLLaMA
Qwen3.8-Flash-Next MTP merged in ik_llama.cpp (integrated head or separate -md file)... 45 → 90 tok/s on a 5090 + 128GB, works down to a 12GB 4070