Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra
This story is from 2026-09-14. It is preserved in the archive; the latest stories are on the live feed.
Model File : https://huggingface.co/unsloth/Qwen3.8-Flash-Next-GGUF Q4_K_XL llama.cpp configuration through llama-swap: -c 256000 --jinja --temp 1 --top-p 0.95 --top-k 20 --min-p 0.00 --presence-penalty 0 --repeat-penalty 1.0 Testing by: llama-benchy Results: | model | test | t/s | peak t/s | ttfr…
Read the full story at r/LocalLLaMA ↗
Timeline · 2 reports
- 2026-09-15 21:33 · r/LocalLLM
Released: Qwen3.8 Flash-Next REAP-384 oQ4e with the full native 512-expert MTP embedded - 2026-09-14 04:30 · r/LocalLLaMA
Data point: Qwen3.8-Flash-Next PP/TG speed on M3 Ultra