Alignment Forecasting: Predicting Misalignment From Training Data
arXiv:2609.35805v1 Announce Type: new Abstract: Training a language model on data with a narrow flaw can sometimes make the model broadly misaligned. Inspecting the data at face value often does not settle whether it will emerge, and today it is caught only after training, by auditing the resulting…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-01 04:00 · arXiv cs.CL
Alignment Forecasting: Predicting Misalignment From Training Data