Composition, Not Conversation: VLMs Lose the Scene, Not the Thread
arXiv:2609.38368v1 Announce Type: new Abstract: Vision-language models (VLMs) increasingly reason over visual evidence that is cropped, segmented, retrieved, or revealed over time. Yet most VQA benchmarks present the complete image and question at once. We ask what models lose when the same informa…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv cs.CV
Composition, Not Conversation: VLMs Lose the Scene, Not the Thread