ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.03355v1 Announce Type: new Abstract: Logit-based knowledge distillation for autoregressive language models usually aligns teacher and student next-token distributions over the entire vocabulary. However, this global objective overlooks relative preferences among likely token alternatives…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-04 04:00 · arXiv stat.ML
ALRA: Adaptive Local Relational Alignment for Logit-Based Pre-training Distillation of Autoregressive Language Models