Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it
Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks. Key findings: Cross-architect…
Read the full story at r/machinelearningnews ↗
Timeline · 2 reports
- 2026-10-02 18:01 · r/deeplearning
Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it - 2026-10-02 18:01 · r/machinelearningnews
Direct weight surgery from Qwen-4B to 0.8B on an 8GB RX 580: why editing all layers breaks everything, and how 4 anchor blocks fixed it