Learning Style, Forgetting Semantics: A Case Study of SFT and RFT on Classification Tasks
arXiv:2610.02437v1 Announce Type: new Abstract: Why does supervised fine-tuning (SFT) lead to more forgetting than reinforcement fine-tuning (RFT), even when all teacher demonstrations are semantically correct? We study this question on classification tasks where tokens within each semantic class e…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-10-05 04:00 · arXiv stat.ML
Learning Style, Forgetting Semantics: A Case Study of SFT and RFT on Classification Tasks