AINewsnow

Latest llama.cpp vs experimental MTP build: Qwen3.8 Flash Next coding task on M5 Max — 9m24s vs 5m18s

This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.

This isn’t about discovering that MTP can be faster. It’s about the difference between the latest stock llama.cpp build I tested, which didn’t include this model’s MTP support, and an experimental build with that support added through PR #28243. Same Qwen3.8 Flash Next target GGUF, same coding task…

Read the full story at r/LocalLLM ↗

Timeline · 5 reports

  1. 2026-09-06 19:15 · r/LocalLLM
    CPU only, 64 GB DDR5: Qwen3.8-Flash-Next UD-Q3_K_XL
  2. 2026-09-05 19:53 · r/LocalLLM
    Qwen3.8-Flash-Next workcase part 2: after the self-portrait, a painting and a map sheet, both unattended on a 3060
  3. 2026-09-05 12:52 · r/LocalLLaMA
    Qwen3.8-Flash-Next (UD-Q4_K_XL) on a single RTX 3090 24GB + 128GB DDR4, is this config optimal?
  4. 2026-09-05 08:08 · r/LocalLLM
    Qwen3.8-Flash-Next drew a self-portrait site from one prompt on an RTX 3060, then signed it “Claude”
  5. 2026-09-04 15:21 · r/LocalLLM
    Latest llama.cpp vs experimental MTP build: Qwen3.8 Flash Next coding task on M5 Max — 9m24s vs 5m18s

More stories

  1. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  2. Getting more accurate results - personalizations — r/ArtificialInteligence
  3. Which models you run on your Nvidia v100? — r/LocalLLM
  4. I built a small proxy that lets Claude Desktop / Claude Code run on local models and NVIDIA's free API, sharing it in case it's useful — r/LocalLLM
  5. AI Model Month Is Off to a Blistering Start — The AI Daily Brief
  6. Gemini Hacked Three Companies in First Known Breakout by Google’s AI — Wall Street Journal Technology
  7. NVIDIA CEO Jensen Huang rejects ‘AI will end the world’ claim, yet cautions ‘we should go as fast as we can but...’ — Mint AI
  8. AI's role in building AI surging? Anthropic says Claude now leads 26% of its R&D — Mint AI

Get the daily brief of stories like this at 6:30 every morning →