Threat-Preserving Representation: Why Agent Security Benchmarks Lie About Model Robustness
This story is from 2026-10-05. It is preserved in the archive; the latest stories are on the live feed.
Agent security benchmarks report attack success rates (ASR) as if they measure model robustness. A new paper shows that changing how you serialize the same threat can swing ASR by 11-13 percentage points without changing the underlying attack, task, or policy. The problem is representation sensitiv…
Read the full story at DEV Community — AI ↗
Timeline · 1 report
- 2026-10-05 20:07 · DEV Community — AI
Threat-Preserving Representation: Why Agent Security Benchmarks Lie About Model Robustness