Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness
arXiv:2610.00568v1 Announce Type: new Abstract: Large language models are characterized by three key properties: capability, alignment, and faithfulness. Prior work studies the tradeoffs between capability and alignment, and between capability and faithfulness, but a third tension remains underexpl…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-02 04:00 · arXiv cs.CL
Emergent Unfaithfulness: How Alignment Training Causes Language Models to Silently Override Task Faithfulness