Stanford's Terminal-Bench-Science Exposes GPT-6 Astra Leading at 63% on Real Science Tasks
A new 70-task agentic benchmark tests frontier models on real scientific research workflows, with top scores stuck in the low 60s.
Read the full story at AlphaSignal ↗
Timeline · 1 report
- 2026-09-24 23:30 · AlphaSignal
Stanford's Terminal-Bench-Science Exposes GPT-6 Astra Leading at 63% on Real Science Tasks