Spomin - Live KV cache compaction (Experimental for Qwen)
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
I’ve been building "Spomin", a router that replaces context with summaries directly in the KV cache. The goal is to keep long running sessions going without repeatedly stopping for full compaction and reprocessing the context that remains. It uses my llama.cpp fork - https://github.com/alekk89/llam…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-11 08:55 · r/LocalLLaMA
Spomin - Live KV cache compaction (Experimental for Qwen)