AINewsnow

VS Code Copilot + llama-server --parallel 2: concurrent requests cause full re-prefill

27B GGUF, --parallel 2 , --kv-unified , --cache-idle-slots . Copilot sends a background summarization request concurrently with the main chat request. Same ~119k token prompt, but the summarization gets cached_tokens = 0 → ~207s full re-prefill instead of ~2s. Root cause: slot selection skips busy…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-10-11 07:19 · r/LocalLLM
    VS Code Copilot + llama-server --parallel 2: concurrent requests cause full re-prefill

More stories

  1. Qwen 3.6 35B A3B: 131K context + vision on 6GB VRAM — r/LocalLLaMA
  2. Tested Mellum2.1-12B-A2.5B on PI Coding Agent - surprisingly usable, but not great at one-shot projects — r/LocalLLaMA
  3. Running the uncensored Qwen3.8-27B (HauhauCS) on a 4090 at 262K context and ~130 tok/s — r/LocalLLaMA
  4. Qwen3.8-Flash-Next (125B) at ~100 tok/s on an M5 Ultra Mac Studio with llama.cpp — r/LocalLLM
  5. Success with Qwen3.8 27B GSQ-RCO-IQ3_S on 16GB VRAM — r/LocalLLM
  6. Created laya : Now Introducing a new 800 Million Param physics-based typed decision model with 73k context and image support — r/LocalLLaMA
  7. Pancho The Llama (Pt I) - Mean People Suck. — r/comfyui
  8. local semantic file search for Linux (Rust, llama.cpp, EmbeddingGemma 2) — r/LocalLLaMA

Get the daily brief of stories like this at 6:30 every morning →