Do Frontier Models Seek Safety Evidence Before Acting?
This story is from 2026-09-17. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.17865v1 Announce Type: new Abstract: Frontier models are often evaluated on how they respond to safety information once it is already in context. We study an earlier decision point: whether models choose to acquire safety-relevant evidence before acting. We introduce SAFE, a controlled b…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-17 04:00 · arXiv cs.AI
Do Frontier Models Seek Safety Evidence Before Acting?