Nemotron-3.5-Lightning at 11.77 GiB, a 16 GB option for a model that didn't have one
This story is from 2026-08-29. It is preserved in the archive; the latest stories are on the live feed.
TL;DR: Every public low-bit GGUF of this model is secretly ~4.70 bpw. Shim the rows to 256 and it becomes a real 3.07 bpw / 11.77 GiB file that runs 262K context on 16GB. Needs patched llama.cpp — not LM Studio or Ollama. In the AtomicChat HuggingFace repo it says "There is currently no good 16 GB…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-08-29 23:27 · r/LocalLLaMA
Nemotron-3.5-Lightning at 11.77 GiB, a 16 GB option for a model that didn't have one