Query-aligned video frame selection for long video understanding
arXiv:2609.31668v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) process multimodal inputs by converting text, images, and videos into token sequences that are subsequently processed by a backbone language model. While MLLMs have achieved excellent performance in understandi…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-29 04:00 · arXiv cs.CV
Query-aligned video frame selection for long video understanding