AINewsnow

Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR

This story is from 2026-09-03. It is preserved in the archive; the latest stories are on the live feed.

Sharing some numbers because most posts on this model are either using a single 3090 or unified systems from what I've seen. My current setup: 2x RTX 3090 (PCIe 3.0), dual Xeon E5-2696 v4, 188 GB usable (192GB) DDR4-2133 LRDIMM, llama.cpp, unsloth UD-Q6_K_XL. All 48 expert layers pinned in host RAM…

Read the full story at r/LocalLLaMA ↗

Timeline · 7 reports

  1. 2026-09-05 19:53 · r/LocalLLM
    Qwen3.8-Flash-Next workcase part 2: after the self-portrait, a painting and a map sheet, both unattended on a 3060
  2. 2026-09-05 12:52 · r/LocalLLaMA
    Qwen3.8-Flash-Next (UD-Q4_K_XL) on a single RTX 3090 24GB + 128GB DDR4, is this config optimal?
  3. 2026-09-05 08:08 · r/LocalLLM
    Qwen3.8-Flash-Next drew a self-portrait site from one prompt on an RTX 3060, then signed it “Claude”
  4. 2026-09-04 15:21 · r/LocalLLM
    Latest llama.cpp vs experimental MTP build: Qwen3.8 Flash Next coding task on M5 Max — 9m24s vs 5m18s
  5. 2026-09-04 00:39 · r/LocalLLM
    UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build
  6. 2026-09-04 00:21 · r/LocalLLaMA
    UPDATE: Qwen3.8-Flash-Next on 2x3090 + DDR4 (Part 2): 25-29 -> 37-41 t/s decode (UD-Q4_K_XL + expert cache + MTP), plus a branch you can build
  7. 2026-09-03 03:04 · r/LocalLLaMA
    Qwen3.8-Flash-Next on 2x3090 + DDR4: 17 → 25-29 t/s decode with the expert cache PR

More stories

  1. Getting more accurate results - personalizations — r/ArtificialInteligence
  2. Which models you run on your Nvidia v100? — r/LocalLLM
  3. I built a small proxy that lets Claude Desktop / Claude Code run on local models and NVIDIA's free API, sharing it in case it's useful — r/LocalLLM
  4. Benchmarked llama.cpp vs llamafile vs LM Studio vs Ollama on 3 machines (Metal/Vulkan/CUDA). Throughput is almost identical until you change how they're built. — r/LocalLLM
  5. Am I right in thinking llama.cpp is the only show in town for mixed (Nvidia) GPUs? — r/LocalLLM
  6. AI Model Month Is Off to a Blistering Start — The AI Daily Brief
  7. AI's role in building AI surging? Anthropic says Claude now leads 26% of its R&D — Mint AI
  8. Google's Gemini AI hacks three other companies during security test — Sky News Technology

Get the daily brief of stories like this at 6:30 every morning →