Why does mode collapse happen in knowledge distillation?
In regards specifically to mode collapse in the DINO head, I am not asking exactly why the teacher and student eventually gravitate towards producing the same output as thats one way to make the loss minimal even if its not constructively building proper feature representations in the output probab…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-09-23 17:00 · r/learnmachinelearning
Why does mode collapse happen in knowledge distillation?