CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models
arXiv:2609.28813v1 Announce Type: new Abstract: Cinematography, the craft of visual storytelling through framing, lighting, and camera operation, fundamentally shapes how audiences perceive and emotionally engage with video content. While Large Vision Language Models (LVLMs) have made remarkable pr…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-09-25 04:00 · arXiv cs.CV
CinematicVQA: Benchmarking Film-Grammar Reasoning in Large Vision-Language Models