Distillation Is Quietly Erasing Your LoRA Preference Tuning
TL;DR — Post-training pipelines chain LoRA fine-tuning, preference optimization (DPO/RLHF), and distillation as if they're interchangeable quality knobs. They aren't. Low-rank adapters often lack the capacity to represent sharp preference distinctions, and distilling on teacher outputs alone throws…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-28 13:16 · DEV Community — Machine Learning
Distillation Is Quietly Erasing Your LoRA Preference Tuning