Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.
This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.
Hey guys, After my CPU-only to 96GB VRAM test, I tested Qwen3.8-Flash-Next across llama.cpp, SGLang and FreeToken on the same workstation. This time I wanted to see what changes when you keep the hardware and model family fixed, but change the engine, weight format and memory placement. I also test…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-09 12:36 · r/LocalLLaMA
What settings do you use for running Qwen3.8-Flash-Next in llama.cpp? - 2026-09-08 19:26 · r/LocalLLaMA
Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines.