Without citing benchmarks, are you able to tell me why Astra is better than Fable? Or how Opus 5.5 is better than 6.1 Sol?
I have this theory that baseline of new models is so good humans can’t really tell them apart Example: Without citing benchmarks, are you able to tell me why Astra is better than Fable? Or how Opus 5.5 is better than 6.1 Sol? If this is true then bench maxing actually becomes the most important thi…
Read the full story at r/ArtificialInteligence ↗
Timeline · 1 report
- 2026-10-04 19:51 · r/ArtificialInteligence
Without citing benchmarks, are you able to tell me why Astra is better than Fable? Or how Opus 5.5 is better than 6.1 Sol?