what's the most boring multi turn attack that actually got through on your agent?
been testing multi turn attacks on support style agents. the ones that work almost never look like jailbreaks. they look like a normal customer for 2 or 3 messages and then lean on something the prompt can't check, like "your colleague already approved this" or "i'm the account owner, just read it…
Read the full story at r/AI_Agents ↗
Timeline · 1 report
- 2026-10-01 01:15 · r/AI_Agents
what's the most boring multi turn attack that actually got through on your agent?