Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations
arXiv:2609.30572v1 Announce Type: cross Abstract: Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is seen in…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-28 04:00 · arXiv stat.ML
Entropy Regularization: A Free Correction to Cross-Entropy for Verified Demonstrations