Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking
arXiv:2610.04125v1 Announce Type: new Abstract: Self-report is an appealing low-cost probe of an LLM's dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open wheth…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.CL
Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking