Aligned Data Can Induce Misalignment via Context Confusion
arXiv:2609.38379v1 Announce Type: new Abstract: Large language models (LLMs) are frequently updated for various use cases, where filtering out misaligned training samples is a common practice for preventing post-update misalignment. However, alignment is inherently context-dependent: a recommendati…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv cs.AI
Aligned Data Can Induce Misalignment via Context Confusion