I mapped 100 papers on AI agent security. Most prompt-injection defenses break under adaptive attacks, here's what holds up on paper
Most prompt-injection defenses that researchers have tested against adaptive attackers failed. The designs with guarantees stop relying on the model and restrict what the agent can do after it reads untrusted text. I build agents with tool access, so I wanted to know which defenses survive an attac…
Read the full story at r/AI_Agents ↗