Qwen3.8-Flash-Next on MLX-serve, 1m context is released!
This story is from 2026-09-09. It is preserved in the archive; the latest stories are on the live feed.
Hi, I'm the co-creator of this Qwen3.8-Flash-Next engine support in MLX-serve. I've been tuning this one to run both fast, efficient and correct up 1m context using kv cache 8 bits in M5 Max 128GB. Qwen is working well at very long context as showed in the video (a snapshot at ~760k context), I let…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-10 16:14 · r/LocalLLM
Qwen3.8-Flash-Next-NVFP4 vs DeepSeek-v4-Flash-0731-FP8 - 2026-09-09 01:34 · r/LocalLLaMA
Qwen3.8-Flash-Next on MLX-serve, 1m context is released!