AINewsnow

Adaptive KV-Cache Streaming V2: Full Context MTP

Hello all, it’s me again. Just a week after my previous post, I started working on a better implementation of Adaptive KV Streaming. Now I’m sharing my second implementation: V2, with full-context MTP. https://github.com/RaymondHuang210129/llama.cpp-adaptive-kv-streaming/tree/feature/adaptive-kv-st…

Read the full story at r/LocalLLM ↗

Timeline · 1 report

  1. 2026-09-29 04:02 · r/LocalLLM
    Adaptive KV-Cache Streaming V2: Full Context MTP

More stories

  1. Best current Qwen Flash Next Q4-ish? + worth using? — r/LocalLLaMA
  2. Qwen3.8-Flash-Next on 12GB VRAM - 65 tokens per second — r/LocalLLM
  3. Full Qwen3.8-27B on one RTX 4090, native Windows: 5,000 tok/s prefill (1.8x llama.cpp) and up to 289 tok/s decode. Ternary Bonsai 27B reaches 532. — r/LocalLLM
  4. Llama.cpp and new model releases ... is Great is the enemy of Good in the LLM world? — r/LocalLLM
  5. Qwen 3.8 27B on a 3090 with a Sonnet 5.5 as a planner: 2.7x cheaper, real numbers — r/LocalLLM
  6. Qwen 3.8 is a workhorse — r/LocalLLaMA
  7. 95+ TPS through 100K generated for qwen3.8 27b, 262K ctx, on a single 3090 — r/LocalLLaMA
  8. FIXED: HTTP 400: Failed to load model "[Specific_Model_Name_In_Use_HERE]". Error: Engine protocol runtime llama-server for [your_chat_session_number_HERE] exited before becoming healthy. exitCode=1, signal=null — r/LocalLLM

Get the daily brief of stories like this at 6:30 every morning →