I benchmarked 10 open-source prompt-injection detectors. The best caught 51%. How are you actually defending your agents?
If your agent reads anything it didn't write — emails, docs, web pages, tool outputs — someone can hide instructions in there and it may follow them. I wanted real numbers on how well the popular defenses hold up, so I ran 10 open-source injection detectors against 629 realistic attacks buried in n…
Read the full story at r/AI_Agents ↗