Adaptive Multi-Value Control in LLMs via Causal Activation Steering
arXiv:2609.30405v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modify…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-28 04:00 · arXiv cs.LG
Adaptive Multi-Value Control in LLMs via Causal Activation Steering