Mask-Guided KV Cache Eviction in Block Diffusion Language Models
arXiv:2610.06996v1 Announce Type: new Abstract: Block diffusion language models keep a large key-value (KV) cache throughout generation and attend to it at every denoising step, limiting both memory capacity and generation speed. Reducing these costs requires deciding which past tokens to use for d…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv cs.LG
Mask-Guided KV Cache Eviction in Block Diffusion Language Models