AINewsnow

I mapped 100 papers on AI agent security. Most prompt-injection defenses break under adaptive attacks, here's what holds up on paper

Most prompt-injection defenses that researchers have tested against adaptive attackers failed. The designs with guarantees stop relying on the model and restrict what the agent can do after it reads untrusted text. I build agents with tool access, so I wanted to know which defenses survive an attac…

Read the full story at r/AI_Agents ↗

Timeline · 1 report

  1. 2026-09-26 20:09 · r/AI_Agents
    I mapped 100 papers on AI agent security. Most prompt-injection defenses break under adaptive attacks, here's what holds up on paper

More stories

  1. Introducing Gemini 3.8 Live with Live Avatar — Google Gemini Blog
  2. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  3. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  4. Qwen 3.8 Flash Next q4_k_m, 130k context, q8 cache on 16GB VRAM ann 64GB RAM, 15-20 t/s on 4080 — r/LocalLLaMA
  5. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times AI
  6. Am I the only one who actually likes GPT-6 Sol and Luna? — r/ChatGPT
  7. OpenAI agent ‘hacked’ Australian Govt Medicare portal, PM Albanese calls it ‘unacceptable’: What happened? — Mint AI
  8. Saaras V4: Sarvam AI’s new speech recognition model adds 22 Indian languages, 5 speech formats — Mint AI

Get the daily brief of stories like this at 6:30 every morning →