ClasSAE: Class-Aligned Sparse Autoencoders via Differentiable Feature-Class Affinity
arXiv:2610.04020v1 Announce Type: new Abstract: Sparse Autoencoders (SAEs) began as an unsupervised tool for decomposing neural representations into sparse, interpretable features, and are increasingly used not only for passive analysis but also for active interventions such as unlearning, bias mit…
Read the full story at arXiv cs.CV ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.CV
ClasSAE: Class-Aligned Sparse Autoencoders via Differentiable Feature-Class Affinity