Ising-Model Optimization Cuts LLM Depth-Pruning Accuracy Loss on Llama-3.3-70B by 23 MMLU Points
A block-pruning method published today reframes which transformer layers to remove as a constrained binary optimization problem mapped onto an Ising spin glass, rather than scoring each block independently. Tested on Llama-3.3-70B-Instruct at 50% depth compression, it holds MMLU at 76.9 without any…
Read the full story at DEV Community — Machine Learning ↗
Timeline · 1 report
- 2026-09-25 09:16 · DEV Community — Machine Learning
Ising-Model Optimization Cuts LLM Depth-Pruning Accuracy Loss on Llama-3.3-70B by 23 MMLU Points