128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build.
I work on Atomic Agent, an open-source agent built for local models first, so weigh this accordingly. We just shipped the first desktop build. This post is about what's under it, because most agents are designed around a frontier cloud model and then "support" local ones, and we tried to go the oth…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-09 00:41 · r/LocalLLM
128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build.