Is there any relation between model merging and distillation(KD,OPD,...)?
Has there been any work exploring the connection between model merging and distillation-based methods(KD,OPD...)? I believe there should be a connection between methods based on parameter space and methods based on output distribution, which may help people understand LLMs.
Read the full story at r/deeplearning ↗
Timeline · 1 report
- 2026-10-07 16:01 · r/deeplearning
Is there any relation between model merging and distillation(KD,OPD,...)?