Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs
arXiv:2609.25635v1 Announce Type: new Abstract: Large Vision-Language Models (LVLMs) incur high computational costs from redundant visual tokens. Although training-free attention-based multi-layer pruning in the vision encoder stage has been explored as an effective strategy, we find that pruning i…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-23 04:00 · arXiv cs.CV
Shallow to Deep: Aligning Token Pruning with Stage-wise Roles in LVLMs