AINewsnow

Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080

Thought it's about time to share after testing for a week. You need four things most people miss: the right quant, the right model, the right branch, and the right cache flags. https://github.com/dtm-beep/qwen38-flash-next-mtp-16gb TLDR: AtomicChat AD-4.27bpw Q4_K_M target + the shared Unsloth MTP…

Read the full story at r/LocalLLaMA ↗

Timeline · 6 reports

  1. 2026-09-26 13:02 · r/LocalLLM
    We have implanted 100 facts into the engram table of Qwen 3.8 Flash Next, and we have now created a website to explain it.
  2. 2026-09-26 08:52 · r/LocalLLaMA
    Best current Qwen Flash Next Q4-ish? + worth using?
  3. 2026-09-26 03:42 · r/LocalLLM
    PSA: Dual 3090 - Qwen Flash Next - 80tps/2k+ prefill
  4. 2026-09-25 06:32 · r/LocalLLaMA
    Did anyone do a full bench of e.g. Qwen Flash Next IQ4 and Qwen 27b FP8? Here are some
  5. 2026-09-24 23:28 · r/LocalLLaMA
    Is Qwen Flash Next at like Q2 better than 27B at Q4?
  6. 2026-09-23 23:46 · r/LocalLLaMA
    Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080

More stories

  1. Google invented the Transformer, so why does using Gemini still feel like a chore? — r/ArtificialInteligence
  2. Jetson Thor — r/LocalLLM
  3. viggle-turbo isn't just faster - for most prompts, it's just as good — r/StableDiffusion
  4. Qwen-Image 2.1 prompt enhancer in ComfyUI: 1.4–1.7× faster with MTP, plus uncensored (heretic) versions — r/StableDiffusion
  5. Viggle/Qwen-Image-2.1-viggle-turbo · v0.2.1 update. — r/StableDiffusion
  6. How do I run comfy inside a venv — r/comfyui
  7. Run Qwen 3.8 27b on the Apple Neural Engine at 7 watts on a Mac — r/LocalLLM
  8. New few-step LoRA adapters for Qwen-Image-2.1, ~5x faster — r/StableDiffusion

Get the daily brief of stories like this at 6:30 every morning →