AINewsnow

Got Qwen3.8-Next-Flash ngram SSD offload working in llama.cpp!

This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.

TL;DR - Save 25% RAM by SSD offloading ngrams with --mmap just by fixing the layout of the Unsloth quant. Tested working on Mac. Thread deleted in LocalLLaMa due to their dumb megathread idea, so reposting here. So, one of the things that excited me about the new Qwen4 arch is the ngram table, expo…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-08-26 20:37 · r/LocalLLM
    Got Qwen3.8-Next-Flash ngram SSD offload working in llama.cpp!

More stories

  1. M2 Mac ultra128gb Qwen flash next — r/LocalLLM
  2. Qwen3.8-Flash-Next-Heretic2-IQ4XS on Halogen Flash Server vs llama-server on Strix Halo: 2.3-7.7x prefill speedup with half the VRAM (+ vision works on BYO GGUF) — r/LocalLLM
  3. Multi-hour llama.cpp optimization experiments on Qwen MoE models, patches, benchmarks, and reproduction guides — r/LocalLLM
  4. You can use any LLM just like JEV — r/LocalLLaMA
  5. The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks — r/LocalLLaMA
  6. CUDA: enable sparse fa for qwen4 by am17an · Pull Request #28770 · ggml-org/llama.cpp — r/LocalLLaMA
  7. focus-llama: a llama.cpp fork implementing Declarative Attention (arXiv:2609.02737) — r/LocalLLaMA
  8. I ran Opencode and PI against the same local model on 3 identical projects, same prompts, same hardware... — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →