SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning
arXiv:2610.11132v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) equips large language models (LLMs) with specialized capabilities, but often comes at the cost of forgetting the general capabilities of their parent models (i.e., the pretrained models before fine-tuning). This trade-off…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-09 04:00 · arXiv cs.CL
SFT-as-Context Mitigates Forgetting in Supervised Fine-Tuning