Released: Qwen3.8 Flash-Next REAP-384 oQ4e with the full native 512-expert MTP embedded
This story is from 2026-09-15. It is preserved in the archive; the latest stories are on the live feed.
I’ve been working on a slightly different approach to Qwen3.8 Flash-Next REAP builds and finally pushed the model to Hugging Face: https://huggingface.co/mensaprodigy/Qwen3.8-Flash-Next-REAP-384-mlx-mtp The basic idea: Prune the expensive 48-layer target model, but leave the native MTP predictor in…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-15 21:33 · r/LocalLLM
Released: Qwen3.8 Flash-Next REAP-384 oQ4e with the full native 512-expert MTP embedded