A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation
arXiv:2609.22109v1 Announce Type: new Abstract: Selective on-policy distillation trains a student only at the token positions a selector scores highest, and the literature compares selectors under a single shared learning rate--a control chosen to be neutral. We show it is not. Under LoRA on GSM8K…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-22 04:00 · arXiv cs.LG
A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation