VS Code Copilot + llama-server --parallel 2: concurrent requests cause full re-prefill
27B GGUF, --parallel 2 , --kv-unified , --cache-idle-slots . Copilot sends a background summarization request concurrently with the main chat request. Same ~119k token prompt, but the summarization gets cached_tokens = 0 → ~207s full re-prefill instead of ~2s. Root cause: slot selection skips busy…
Read the full story at r/LocalLLM ↗
Timeline · 1 report
- 2026-10-11 07:19 · r/LocalLLM
VS Code Copilot + llama-server --parallel 2: concurrent requests cause full re-prefill