HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization
This story is from 2026-09-30. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.35800v1 Announce Type: new Abstract: Low-bit key-value (KV) cache quantization saves storage but can sharply degrade vision-language model (VLM) accuracy. We introduce HeadGuard, a composable head-protection method that augments a base KV-cache quantizer with a fixed high-precision mask.…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-30 04:00 · arXiv cs.LG
HeadGuard: Selective Head Protection for Low-Bit VLM KV-Cache Quantization