Fixed-weight models are adversarially vulnerable: hence misaligned
This post argues that fixed-weight models (at least as we understand them today) will a) always be vulnerable to adversarial examples in their concept-spaces, and b) hence will be misaligned, under sufficient optimisation pressure . Boundaries in concept space To serve any purpose whatsoever, an AI…
Read the full story at Alignment Forum ↗
Timeline · 1 report
- 2026-09-28 20:08 · Alignment Forum
Fixed-weight models are adversarially vulnerable: hence misaligned