AINewsnow

128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build.

I work on Atomic Agent, an open-source agent built for local models first, so weigh this accordingly. We just shipped the first desktop build. This post is about what's under it, because most agents are designed around a frontier cloud model and then "support" local ones, and we tried to go the oth…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-09 00:41 · r/LocalLLM
    128K context on Qwen 3.5 4B in 800 MB instead of 4 GB: what we changed in our llama.cpp build.

More stories

  1. RPC: add `-sm tensor` by am17an · Pull Request #26610 · ggml-org/llama.cpp — r/LocalLLaMA
  2. Running a local server with Gemma 4 26b a4b on laptop rtx 4050 + 16gb ram dd5 and llama.cpp — r/LocalLLM
  3. Looking for developer-friendly inference providers who give you enough API credits to experiment [D] — r/MachineLearning
  4. unsloth/Qwen3.8-Flash-Next-GGUF is being updated — r/LocalLLaMA
  5. NInfer6000 - Qwen 3.8 Flash Next @ 400 tg/s & 13K pp/s — r/LocalLLaMA
  6. Story time: Qwen3.8-Flash-Next on my Strix Halo laptop vs Claude Opus 5.5 on the same feature — r/LocalLLaMA
  7. Java vllm-like framwork claims 90% of perfomance of llama.cpp on local inference on NVIDIA GPUs by compiling Java to CUDA and cuTile — r/LocalLLM
  8. Worth moving on from Qwen3.6 35B A3B UD on a gaming PC? — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →