BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents
This story is from 2026-09-16. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.16305v1 Announce Type: new Abstract: Large language model (LLM) agents increasingly operate over long-horizon interactions involving tool use, persistent state, evolving authorization, and external environment feedback. In such settings, safety failures may emerge only after multiple tur…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-16 04:00 · arXiv cs.AI
BLINDSPOT: A Benchmark for Safety and Refusal Calibration in Long-Horizon Tool-Using Agents