Stanford's MAttr Tops AI Interpretability Benchmark by Nearly 3x
A new interpretability method learns which internal components matter for a behavior, and pinpoints just 1% of Llama 3.1 weights driving refusals.
Read the full story at AlphaSignal ↗
Timeline · 1 report
- 2026-09-24 06:01 · AlphaSignal
Stanford's MAttr Tops AI Interpretability Benchmark by Nearly 3x