I ran 100 Terminal-Bench 2.1 slots on Luna 5.6 and Luna 6. The Luna 6 results still look like a joke
I've run this for the first time 2 days ago. I was shocked by the results, so I wanted to check whether the big gap I saw between GPT-5.6 Luna and GPT-6 Luna was just a bad run, so I ran the comparison again today. The tasks came from Terminal-Bench 2.1 on Harbor . I selected the 100 shortest trial…
Read the full story at r/OpenAI ↗
Timeline · 1 report
- 2026-09-25 19:14 · r/OpenAI
I ran 100 Terminal-Bench 2.1 slots on Luna 5.6 and Luna 6. The Luna 6 results still look like a joke