PSA: llama.cpp -cram should be increased for agentic workflows (default is 8192)
Just a quick PSA. llama.cpp does have prompt caching. if you are running large context lengths and have long multiturn projects, increasing -cram can provide you with massive speedups. There is a point where context lengths can get so large that 8192mb is not enough and the whole context needs to b…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-24 15:44 · r/LocalLLaMA
PSA: llama.cpp -cram should be increased for agentic workflows (default is 8192)