exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup!
This story is from 2026-09-07. It is preserved in the archive; the latest stories are on the live feed.
I got 2x 20GB RTX 3080s + 128GB of DDR4 2666hz RAM (only 4 of 6 channels populated) + a Xeon 6148 I've always been a llama.cpp person and I've been running Unsloth's Q4_K_XL quant of Qwen 3.8 Flash Next at ~270tps prefill and ~13tps decode (starts off close to 20 and falls down to 13 with growing c…
Read the full story at r/LocalLLaMA ↗
Timeline · 4 reports
- 2026-09-09 12:36 · r/LocalLLaMA
What settings do you use for running Qwen3.8-Flash-Next in llama.cpp? - 2026-09-08 19:26 · r/LocalLLaMA
Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines. - 2026-09-08 00:27 · r/LocalLLM
An optimized llama.cpp for people wanting to run Qwen 3.8 Flash Next on two Volta v100 32gbs - 2026-09-07 19:16 · r/LocalLLaMA
exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup!