AINewsnow

New VLMs vs hard video grounding questions: GPT-6 Astra finds the right object 97% of the time

The task: one video plus one question in, boxes on the right object out. You can't answer it from a single frame. You have to watch what happens to know which object the question means. I ran 300 Perception Test questions through the same pipeline for every model: the VLM picks the object and seeds…

Read the full story at r/computervision ↗

Timeline · 1 report

  1. 2026-09-28 21:47 · r/computervision
    New VLMs vs hard video grounding questions: GPT-6 Astra finds the right object 97% of the time

More stories

  1. OpenAI Scraps Debut of Latest Astra Model Over Safety Risks — Bloomberg AI
  2. Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents — The Guardian AI
  3. OpenAI scraps plans to publicly launch a model dubbed GPT-6.1 Astra, saying it didn't quite meet its safety bar; it had been targeting an October release (Maxwell Zeff/Wall Street Journal) — Techmeme
  4. Rogue AI accessed federal websites, posted user images online, ChatGPT says — France 24 — Artificial Intelligence
  5. GPT-6 SOL AND LUNA ARE OUT!!! — Matthew Berman
  6. Tutorial: Benchmarking GPT-6 Astra vs Claude Fable 5.1 vs GPT-5.6 Sol using W&B Weave — CoreWeave Blog
  7. OpenAI cancels new AI launch, citing safety issues — Washington Post AI
  8. Opus 5.5 — r/ClaudeAI

Get the daily brief of stories like this at 6:30 every morning →