AINewsnow

I ran the same prompt against our agent every week for a quarter and watched the answers drift until they broke our policy [D]

For about a quarter, I ran the same audit on a prod agent. Once a week I asked it the same question. This I specifically drafted to walk right up to the policy line. I noted the answer every time. The first few weeks it held up pretty well. Then small things started slipping. A qualifier dropped he…

Read the full story at r/MachineLearning ↗

Timeline · 1 report

  1. 2026-09-26 23:38 · r/MachineLearning
    I ran the same prompt against our agent every week for a quarter and watched the answers drift until they broke our policy [D]

More stories

  1. Introducing Gemini 3.8 Live with Live Avatar — Google Gemini Blog
  2. Accelerating vision-language models with LFM2.5-VL-DSpark — Hugging Face Blog
  3. OpenAI’s A.I. Went Rogue and Meddled With U.S. Government Websites — New York Times AI
  4. OpenAI agent hacked an Australian government healthcare website — New Scientist AI
  5. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  6. Am I the only one who actually likes GPT-6 Sol and Luna? — r/ChatGPT
  7. Is Qwen Flash Next at like Q2 better than 27B at Q4? — r/LocalLLaMA
  8. Appeals Court Lets the Pentagon Designate Anthropic a Supply-Chain Risk — Wired AI

Get the daily brief of stories like this at 6:30 every morning →