AINewsnow

Qwen Flash Next MTP work restarted

If you're using MTP with Qwen Flash Next and llama.cpp, you can switch to: quants: https://huggingface.co/ggml-org/Qwen3.8-Flash-Next-GGUF PR: https://github.com/ggml-org/llama.cpp/pull/29761 Please note that this is still wip

Read the full story at r/LocalLLaMA ↗

Timeline · 2 reports

  1. 2026-10-03 01:47 · r/LocalLLM
    Strata is seriously impressive, running Qwen 3.8 Flash Next on hermes at 512k context.
  2. 2026-10-01 05:27 · r/LocalLLaMA
    Qwen Flash Next MTP work restarted

More stories

  1. Qwen4Exp: add MTP by am17an · Pull Request #29761 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Pi extension: Skip reasoning with local Qwen 27B and proceed to answer right now — r/LocalLLaMA
  3. Claude and Grok built me a local monitoring setup for my AI box: two dashboards, one for the machine, one for model training — r/LocalLLM
  4. Who’s the current “king” of local LLMs for you — Qwen, Gemma, Llama, something else? — r/LocalLLM
  5. Which LLM is best for coding/agents if I have dual R9700 GPUs? — r/LocalLLM
  6. Used Opus 5.5 to optimize llama.cpp inference for Swift Qwen 3.8 27B Q6_K on RTX 5090 - decode 143 tok/s prefill 2840 tok/s — r/LocalLLM
  7. microsoft/FrogNano-4B-2609 · Hugging Face — r/LocalLLaMA
  8. Viggle/Qwen-Image-2.1-viggle-turbo · Hugging Face — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →