UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
This is a follow-up to my post from yesterday (17 -> 25-29 t/s with the expert cache PR). Same box: 2x RTX 3090 on PCIe 3.0, dual Xeon E5-2696 v4, 188 GB DDR4-2133 LRDIMM, llama.cpp, full 261k context, f16 KV, all 48 expert layers in host RAM, everything else on the GPUs. Since then I switched quan…
Read the full story at r/LocalLLaMA ↗
Timeline · 7 reports
- 2026-09-06 19:15 · r/LocalLLM
CPU only, 64 GB DDR5: Qwen3.8-Flash-Next UD-Q3_K_XL - 2026-09-05 19:53 · r/LocalLLM
Qwen3.8-Flash-Next workcase part 2: after the self-portrait, a painting and a map sheet, both unattended on a 3060 - 2026-09-05 12:52 · r/LocalLLaMA
Qwen3.8-Flash-Next (UD-Q4_K_XL) on a single RTX 3090 24GB + 128GB DDR4, is this config optimal? - 2026-09-05 08:08 · r/LocalLLM
Qwen3.8-Flash-Next drew a self-portrait site from one prompt on an RTX 3060, then signed it “Claude” - 2026-09-04 15:21 · r/LocalLLM
Latest llama.cpp vs experimental MTP build: Qwen3.8 Flash Next coding task on M5 Max — 9m24s vs 5m18s - 2026-09-04 00:39 · r/LocalLLM
UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build - 2026-09-04 00:21 · r/LocalLLaMA
UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build