Pushing models to their limits: The "Bad Apple" benchmark
I first wanted to do something with GPT-6 Astra, but every time it completely missed what I wanted and constantly tried to cheat, it just couldn't get it right. So, I started from scratch with Opus 5.5, it succeeded surprisingly well, it immediately got what I was looking for, however I had to give…
Read the full story at r/singularity ↗
Timeline · 1 report
- 2026-09-29 02:53 · r/singularity
Pushing models to their limits: The "Bad Apple" benchmark