Qwen3.8-Flash-Next: Time to Update Those Benchmarks
This story is from 2026-08-27. It is preserved in the archive; the latest stories are on the live feed.
specs hardware: M4 Max 128GB Studio inference engine: oMLX & lllama.cpp insights it still very early, so had to disable oMLX K/V caching, qwen4_exp architectureis not yet supported + the obvious n-grams with which the whole 4 bit quant takes ~100G, so pretty tight nevertheless, this is the first mo…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-27 12:34 · r/LocalLLaMA
Qwen3.8-Flash-Next: Time to Update Those Benchmarks