AINewsnow

Qwen 3.8 Flash on 64GB RAM and 8GB VRAM Custom Fork LLama

This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.

This is mostly to explain my experience optimizing the model to run on my machine and maybe getting interest from someone to go and write an actual PR against llamacpp (as I'm absolutely not willing to generalize this code ahah). Short Preface (this post is hand written, no AI here). So when Qwen d…

Read the full story at r/LocalLLM ↗

Timeline · 7 reports

  1. 2026-09-12 11:02 · r/LocalLLaMA
    I am impressed and I owe you one, Qwen 3.8 flash next (vision)!
  2. 2026-09-12 07:02 · r/LocalLLM
    Qwen flash next beats Fable
  3. 2026-09-12 00:01 · r/LocalLLM
    Qwen 3.8-Flash-Next On Mac Mini with 64gb Works
  4. 2026-09-11 22:12 · r/LocalLLM
    Minimal total RAM + VRAM to run Qwen 3.8 flash next with decent speed?
  5. 2026-09-11 14:58 · r/LocalLLM
    Qwen 3.8 Flash Next MTP tested speed
  6. 2026-09-11 10:40 · r/LocalLLaMA
    7900 XTX + 32/64GB RAM for Qwen 3.8 Flash Next?
  7. 2026-09-11 01:28 · r/LocalLLM
    Qwen 3.8 Flash on 64GB RAM and 8GB VRAM Custom Fork LLama

More stories

  1. Deployed Qwen 3.6 35B A3B on a single DGX Spark supporting 12 concurrent users at 262K context. Are there better ways to optimize this? — r/LocalLLM
  2. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  3. Pay $39.99 once to put ChatGPT, Claude, Gemini, and more in a single workspace for life — Mashable AI
  4. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  5. M1 Max 32GB, trying to run Qwen 3.8 27B at decent speeds and context — r/LocalLLaMA
  6. [Guide / Weights] Qwen 3.8 27B on Intel Arc: Why IQ quants crawl at 8 tok/s, why Q4_K outpaces sub-4bpw on Battlemage, and clean RCO GGUFs (16GB & 24GB) — r/LocalLLM
  7. My Version of Jev running locally, playing doom. — r/LocalLLM
  8. Two node BC250 cluster comparison of Qwen3.6 vs Qwen 3.8 — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →