LLM Pareto frontiers split by benchmark category, Sonnet 5.5 pushes last OpenAI models off the frontier
Yeah I know there are a few of these already, but I wanted one split by benchmark category instead of one blended score. Anyway, as I was parsing latest models just noticed Sonnet 5.5 benchmarks pop in and push the last OpenAI models off the frontier. I implemented some feedback I got here previous…
Read the full story at r/ArtificialInteligence ↗
Timeline · 1 report
- 2026-09-28 20:36 · r/ArtificialInteligence
LLM Pareto frontiers split by benchmark category, Sonnet 5.5 pushes last OpenAI models off the frontier