AINewsnow

An optimized llama.cpp for people wanting to run Qwen 3.8 Flash Next on two Volta v100 32gbs

This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.

Disclaimer: this is (mostly) vibed, not gonna pretend otherwise - im just posting in case it helps someone trying this setup. I spent a few days on it and offered it up another guy on here (on request) and he said it gave him some big speedups, and he made a new PR fixing some of my bugs. Provided…

Read the full story at r/LocalLLM ↗

Timeline · 3 reports

  1. 2026-09-09 12:36 · r/LocalLLaMA
    What settings do you use for running Qwen3.8-Flash-Next in llama.cpp?
  2. 2026-09-08 19:26 · r/LocalLLaMA
    Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.
  3. 2026-09-08 00:27 · r/LocalLLM
    An optimized llama.cpp for people wanting to run Qwen 3.8 Flash Next on two Volta v100 32gbs

More stories

  1. Is Qwen 3.8 Flash Next usable on M1 Ultra 64GB ? — r/LocalLLM
  2. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  3. Intel releases OpenVINO 2026.4 — r/LocalLLaMA
  4. M1 Max 32GB, trying to run Qwen 3.8 27B at decent speeds and context — r/LocalLLaMA
  5. [Guide / Weights] Qwen 3.8 27B on Intel Arc: Why IQ quants crawl at 8 tok/s, why Q4_K outpaces sub-4bpw on Battlemage, and clean RCO GGUFs (16GB & 24GB) — r/LocalLLM
  6. My Version of Jev running locally, playing doom. — r/LocalLLM
  7. Two node BC250 cluster comparison of Qwen3.6 vs Qwen 3.8 — r/LocalLLM
  8. dual 7900 xtx - some guy made a pretty optimized fork of lamacpp optimized for this setup Qwen 3.8 Q8 at 82 tokens / seconds decode — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →