GPT-6 Astra makes a massive leap on ZeroBench (an extremely difficult vision benchmark), surpassing the human baseline across all three metrics
pass@5: Scores if at least one of the 5 attempts is correct pass^5: Scores only if all 5 attempts are correct (a reliability metric) "pass@1": Not a true single-attempt pass@1, it's the average score across 5 attempts. https://zerobench.github.io/
Read the full story at r/singularity ↗