AINewsnow

LLM Pareto frontiers split by benchmark category, Sonnet 5.5 pushes last OpenAI models off the frontier

Yeah I know there are a few of these already, but I wanted one split by benchmark category instead of one blended score. Anyway, as I was parsing latest models just noticed Sonnet 5.5 benchmarks pop in and push the last OpenAI models off the frontier. I implemented some feedback I got here previous…

Read the full story at r/ArtificialInteligence ↗

Timeline · 1 report

  1. 2026-09-28 20:36 · r/ArtificialInteligence
    LLM Pareto frontiers split by benchmark category, Sonnet 5.5 pushes last OpenAI models off the frontier

More stories

  1. NVIDIA Open Agent Safety Platform: A Reference for Continuous In-Silicon Agent Monitoring — NVIDIA Technical Blog
  2. How we found 24 Android vulnerabilities using our open source AI security agent — GitHub Blog
  3. OpenAI Scraps Debut of Latest Astra Model Over Safety Risks — Bloomberg AI
  4. Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents — The Guardian AI
  5. OpenAI bots meddled with multiple US government agency sites — BBC Technology
  6. How many times have AI agents gone 'rogue'? OpenAI says review of full scope may take months — Mint AI
  7. Use ChatGPT Work to build your data agent — OpenAI YouTube
  8. OpenAI scraps plans to publicly launch a model dubbed GPT-6.1 Astra, saying it didn't quite meet its safety bar; it had been targeting an October release (Maxwell Zeff/Wall Street Journal) — Techmeme

Get the daily brief of stories like this at 6:30 every morning →