A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass
arXiv:2610.03940v1 Announce Type: new Abstract: During fine-tuning, a language model can assign less probability to previously learned answers even when the current gradient acts to preserve that probability. With momentum, each update also carries gradients computed at earlier model states, and th…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.CL
A Step Towards Forgetting: Optimiser History and the Loss of Answer Mass