Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models
arXiv:2610.04017v1 Announce Type: new Abstract: Localizing behavior to individual components of a language model is a central goal of mechanistic interpretability. However, scoring components one at a time misses context-dependent effects: a primary component can inhibit the activation of a backup,…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.LG
Slaying the Hydra: Interaction-Aware Circuit Discovery in Language Models