Is llama.cpp meant to be slow at long context, even when you aren't using that context?
I am trying out a few fine-tunes of Qwen 3.5 9B @ IQ4_XS @ 131K context and trying to go mostly local (free beats cheap, after all). However, it is much slower than at, say 16K context, even when I am not actually using 131K tokens in the first place. Anyone know why this is? I am using the followi…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-22 11:29 · r/LocalLLaMA
Is llama.cpp meant to be slow at long context, even when you aren't using that context?