Seeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act
arXiv:2609.22588v1 Announce Type: new Abstract: Vision-language models (VLMs) perform strongly on visual question answering benchmarks, yet often make decisions that contradict visual evidence they have already identified correctly. We distinguish perceptual failure, where relevant evidence is not…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv cs.CV
Seeing is not Enough: Vision-Language Models Perceive Evidence but Fail to Act