Got Qwen3.8-Next-Flash ngram SSD offload working in llama.cpp!
This story is from 2026-08-26. It is preserved in the archive; the latest stories are on the live feed.
TL;DR - Save 25% RAM by SSD offloading ngrams with --mmap just by fixing the layout of the Unsloth quant. Tested working on Mac. Thread deleted in LocalLLaMa due to their dumb megathread idea, so reposting here. So, one of the things that excited me about the new Qwen4 arch is the ngram table, expo…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-08-26 20:37 · r/LocalLLM
Got Qwen3.8-Next-Flash ngram SSD offload working in llama.cpp!