Post-Training Leaves Behavioral Shadows on Unrelated Decisions
arXiv:2609.29233v1 Announce Type: new Abstract: We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass throug…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-25 04:00 · arXiv cs.CL
Post-Training Leaves Behavioral Shadows on Unrelated Decisions