An optimized llama.cpp for people wanting to run Qwen 3.8 Flash Next on two Volta v100 32gbs
This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.
Disclaimer: this is (mostly) vibed, not gonna pretend otherwise - im just posting in case it helps someone trying this setup. I spent a few days on it and offered it up another guy on here (on request) and he said it gave him some big speedups, and he made a new PR fixing some of my bugs. Provided…
Read the full story at r/LocalLLM ↗
Timeline · 3 reports
- 2026-09-09 12:36 · r/LocalLLaMA
What settings do you use for running Qwen3.8-Flash-Next in llama.cpp? - 2026-09-08 19:26 · r/LocalLLaMA
Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines. - 2026-09-08 00:27 · r/LocalLLM
An optimized llama.cpp for people wanting to run Qwen 3.8 Flash Next on two Volta v100 32gbs