Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models
arXiv:2610.06945v1 Announce Type: new Abstract: Mechanistic interpretability often relies on the Linear Representation Hypothesis (LRH), which assumes that high-level concepts are encoded as linear directions in activation space. Yet a natural visual concept does not necessarily require a linear vi…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-07 04:00 · arXiv cs.CV
Beyond the Linear Representation Hypothesis: Non-Linear Activation Steering in Text-to-Image Models