AINewsnow

Pushing models to their limits: The "Bad Apple" benchmark

I first wanted to do something with GPT-6 Astra, but every time it completely missed what I wanted and constantly tried to cheat, it just couldn't get it right. So, I started from scratch with Opus 5.5, it succeeded surprisingly well, it immediately got what I was looking for, however I had to give…

Read the full story at r/singularity ↗

Timeline · 1 report

  1. 2026-09-29 02:53 · r/singularity
    Pushing models to their limits: The "Bad Apple" benchmark

More stories

  1. GPT-6 turns my room into Studio Ghibli (on Apple Vision Pro) — r/singularity
  2. OpenAI Scraps Debut of Latest Astra Model Over Safety Risks — Bloomberg AI
  3. Heads of OpenAI and Anthropic called to face Senate inquiry after rogue agent incidents — The Guardian AI
  4. Rogue AI accessed federal websites, posted user images online, ChatGPT says — France 24 — Artificial Intelligence
  5. Use ChatGPT Work to build your data agent — OpenAI YouTube
  6. OpenAI scraps plans to publicly launch a model dubbed GPT-6.1 Astra, saying it didn't quite meet its safety bar; it had been targeting an October release (Maxwell Zeff/Wall Street Journal) — Techmeme
  7. OpenAI postpones release of latest AI model over security concerns as the industry faces new safety pressures — Euronews Next
  8. GPT-6 SOL AND LUNA ARE OUT!!! — Matthew Berman

Get the daily brief of stories like this at 6:30 every morning →