"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It
arXiv:2609.25021v1 Announce Type: new Abstract: Large Language Models (LLMs) tend to add disclaimers like "I'm just an AI" when asked about something related to themselves. The self-reports from such responses are used in debates about AI safety or self-knowledge of the models, yet what drives them…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-23 04:00 · arXiv cs.LG
"As a Language Model...": Chat Template Switches LLM Self-Referential Voice and Activation Steering Reproduces It