Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?
arXiv:2610.04119v1 Announce Type: new Abstract: Mechanistic interpretability usually studies fully trained models, yet the computations that drive a behaviour can change while the model is still learning the task. On the Indirect Object Identification task, a model should continue with the name men…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.CL
Copying Before Suppression: What Drives a Below-Chance Dip During Language Model Training?