New VLMs vs hard video grounding questions: GPT-6 Astra finds the right object 97% of the time
The task: one video plus one question in, boxes on the right object out. You can't answer it from a single frame. You have to watch what happens to know which object the question means. I ran 300 Perception Test questions through the same pipeline for every model: the VLM picks the object and seeds…
Read the full story at r/computervision ↗
Timeline · 1 report
- 2026-09-28 21:47 · r/computervision
New VLMs vs hard video grounding questions: GPT-6 Astra finds the right object 97% of the time