llama.cpp ngram on RAM/SSD?
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
I've been out of the loop for some time. Is there already an official way to offload ngram to RAM or SSD in something like Unsloth Studio? Interested in running Qwen3.8-Flash-Next on 72GB VRAM, but naiive attempts failed because even at Q4 it seems to load God knows what to God knows where. Would a…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-11 14:39 · r/LocalLLaMA
llama.cpp ngram on RAM/SSD?