Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning
This story is from 2026-09-11. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.09707v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) applies a uniform cross-entropy loss to all target tokens, even though different tokens provide unequal learning signals for mathematical reasoning. This uniform treatment can over-sharpen already mastered tokens while amp…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-09-11 04:00 · arXiv cs.AI
Which Tokens Should SFT Actually Learn? A Token-Trimming Perspective on Mathematical Reasoning