Speculative reward hacking in coding agents
I audited thousands of agent rollouts in DeepSWE-1.1. Over 80% contained reasoning about an imagined grader. Yet no grader/verifier is mentioned in prompts nor accessible to the agents. Agents reasoned things like: " Let me look at the problem from the grader's perspective " and referred to " hidde…
Read the full story at r/LocalLLaMA ↗
Timeline · 1 report
- 2026-09-28 23:25 · r/LocalLLaMA
Speculative reward hacking in coding agents