Direct weight surgery from Qwen-4B to 0.8B on an 8GB: why editing all layers breaks everything, and how 4 anchor blocks fixed it
Instead of spending weeks and billions of tokens on standard distillation, we tested an alternative: extracting layer-to-layer hidden state trajectories on a handful of calibration prompts and solving for closed-form weight updates directly in the student's MLP blocks. Key findings: Cross-architect…
Read the full story at r/machinelearningnews ↗
Timeline · 1 report
- 2026-10-05 18:18 · r/machinelearningnews
Direct weight surgery from Qwen-4B to 0.8B on an 8GB: why editing all layers breaks everything, and how 4 anchor blocks fixed it