Asymmetries in Spontaneous and Instructed Deception
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.00180v1 Announce Type: new Abstract: Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-02 04:00 · arXiv cs.AI
Asymmetries in Spontaneous and Instructed Deception