Qwen3.8-27B: Using KV Cache Transplants to Boost Output Quality
Since my last post , I've been thinking about different options for dynamic performance degradation, trying to squeeze as much high-quality inference out of my GPU as I can. Over the weekend I read this really interesting paper: Cache-to-Cache: Direct Semantic Communication Between Large Language M…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-25 20:35 · r/LocalLLaMA
Qwen3.8-27B: Using KV Cache Transplants to Boost Output Quality