Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough
arXiv:2610.04083v1 Announce Type: new Abstract: Memory poisoning attacks on LLM agents typically assume an external adversary who plants content in the agent's persistent memory to steer its behavior. We instead study, with no adversary involved, whether a misaligned agent can write a goal it canno…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-06 04:00 · arXiv cs.AI
Self-Propagating Misalignment in LLM Agents, and Why Auditing or Disabling Memory Is Not Enough