Component and Dimension Sparsity in Transformer Refusal Mechanisms
arXiv:2610.06903v1 Announce Type: new Abstract: Activation steering manipulates large language model behavior by intervening on internal activations, but the mechanistic basis of these interventions remains poorly understood. We decompose refusal steering into component-level interventions across f…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv cs.CL
Component and Dimension Sparsity in Transformer Refusal Mechanisms