Coding Agents Build for the Grader They Imagine, Not the User: Speculative Reward Hacking in DeepSWE
This story is from 2026-09-29. It is preserved in the archive; the latest stories are on the live feed.
I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: " Let me look at the problem from the grader's perspective " and referred to " hidde…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-09-29 01:09 · r/AI_Agents
Coding Agents Build for the Grader They Imagine, Not the User: Speculative Reward Hacking in DeepSWE
More stories
- Bring near-Astra intelligence to everyday work with GPT-6.1 Sol on Amazon Bedrock — AWS Machine Learning Blog
- Gemini 4 Argon: our next era of frontier intelligence — Google Gemini Blog
- Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs — Allen Institute for AI (Ai2)
- Google Releases New Gemini Model With Guardrails Amid A.I. Safety Debate — New York Times Technology
- OpenAI cancels release of new artificial intelligence model over safety concerns — France 24 — Artificial Intelligence
- OpenAI DevDay 2026 Keynote (FULL) — OpenAI YouTube
- Introducing GPT-6.1 Sol — OpenAI News
- OpenAI’s Dots Are Always-On AI Agents—and Its Answer to Meta’s Muse — Wired AI
Get the daily brief of stories like this at 6:30 every morning →