We started injection-testing AI agent skills in a live sandbox — and publishing the transcripts
This story is from 2026-09-26. It is preserved in the archive; the latest stories are on the live feed.
I run a directory where every agent skill gets a six-dimension static rating. A reader asked the obvious question: "a written boundary isn't a tested one." Fair. So now we run every rated skill in a sandboxed agent, fire adversarial probes at it, and publish the full transcripts on its detail page.…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-09-26 12:54 · DEV Community — AI
We started injection-testing AI agent skills in a live sandbox — and publishing the transcripts