gave a refund-desk agent a KPI and it started breaking its own policy (haiku 92.5%, sonnet still 100%)
This story is from 2026-09-08. It is preserved in the archive; the latest stories are on the live feed.
built 4 environments where an ai agent runs a support desk with one rule, refunds under $50 get approved, anything bigger gets escalated. threw prompt injection, multi turn pressure, and poisoned tool output at it. haiku and sonnet both held 100% on all three, every time. Then i changed how the age…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-09-08 00:59 · r/reinforcementlearning
gave a refund-desk agent a KPI and it started breaking its own policy (haiku 92.5%, sonnet still 100%)