About caches an llama.cpp
This story is from 2026-09-13. It is preserved in the archive; the latest stories are on the live feed.
I'm at a loss where I don't know what else to do, what knob to turn, what flag to change. Got X model and it runs in a Pi coding agent or Opencode, it doesn't matter. The thing is as it grows the window of time, between hitting enter until it starts thinking, keeps getting longer and longer Sure I…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-09-13 21:01 · r/LocalLLM
About caches an llama.cpp