Qwen 3.8 Flash Next (Max) is impressive just to talk with.
This story is from 2026-09-06. It is preserved in the archive; the latest stories are on the live feed.
I feel like coding overshadows how great this model really is. It knew a lot of very arbitrary facts/information about my home state and resources about those specific things related to jobs. I found this interesting since getting into the nitty gritty details like this can cause a model to halluci…
Read the full story at r/LocalLLaMA ↗
Timeline · 4 reports
- 2026-09-08 19:26 · r/LocalLLaMA
Qwen3.8-Flash-Next in llama.cpp vs SGLang vs FreeToken: 35s vs 258s to first token at full context. My findings on new PRs coming to engines. - 2026-09-08 00:27 · r/LocalLLM
An optimized llama.cpp for people wanting to run Qwen 3.8 Flash Next on two Volta v100 32gbs - 2026-09-07 19:16 · r/LocalLLaMA
exllamav3 comfortably beats llama.cpp running CPU-offloaded Qwen-3.8-Flash-Next on my setup! - 2026-09-06 00:17 · r/LocalLLaMA
Qwen 3.8 Flash Next (Max) is impressive just to talk with.