Has anyone found a reliable way to scan agent skills for security risks without drowning in false positives? I’m looking for something that can distinguish genuinely unsafe behavior from legitimate tool use, rather than flagging anything that looks suspicious in isolation.
We've spent way too much time looking into ways to check agent skills before using them and I still don't feel great about any of it. AI review can get prompt injected by the thing it's reviewing, scanners flag stuff the skill legitimately needs to do, and so many findings end up needing a manual r…
Read the full story at r/AI_Agents ↗