AINewsnow

I ran 100 Terminal-Bench 2.1 slots on Luna 5.6 and Luna 6. The Luna 6 results still look like a joke

I've run this for the first time 2 days ago. I was shocked by the results, so I wanted to check whether the big gap I saw between GPT-5.6 Luna and GPT-6 Luna was just a bad run, so I ran the comparison again today. The tasks came from Terminal-Bench 2.1 on Harbor . I selected the 100 shortest trial…

Read the full story at r/OpenAI ↗

Timeline · 1 report

  1. 2026-09-25 19:14 · r/OpenAI
    I ran 100 Terminal-Bench 2.1 slots on Luna 5.6 and Luna 6. The Luna 6 results still look like a joke

More stories

  1. GPT‑6 Sol and Luna: Cheaper, but Worse Where It Matters — r/OpenAI
  2. OpenAI agent ‘hacked’ Australian Govt Medicare portal, PM Albanese calls it ‘unacceptable’: What happened? — Mint AI
  3. Meet the Data Agent in ChatGPT Work — OpenAI YouTube
  4. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  5. Need some help — r/AI_Agents
  6. new update? — r/GeminiAI
  7. Question about Wan 3 — r/StableDiffusion
  8. OpenAI rogue agents targeted govt and varsity websites in US, Australia before Hugging Face hack: What we know — Mint AI

Get the daily brief of stories like this at 6:30 every morning →