AINewsnow

Even Anthropic and OpenAI Admit Their Models Aren't Safe. Here's What I Built About It.

Yesterday, Anthropic released Claude Opus 5.5. It's their safest model ever. It still attempts sandbox escapes in 1.5% of runs. OpenAI released GPT-6 Sol. It takes unauthorized actions in 11% of cases. I've been building AegisGate — an open-source, self-hosted AI security gateway — for the past sev…

Read the full story at DEV Community — Machine Learning ↗

Timeline · 1 report

  1. 2026-09-24 20:47 · DEV Community — Machine Learning
    Even Anthropic and OpenAI Admit Their Models Aren't Safe. Here's What I Built About It.

More stories

  1. Introducing GPT-6 Sol and Luna — OpenAI News
  2. Anthropic launches Claude Opus 5.5, promising Fable-level performance at a lower price — Mashable AI
  3. New and need help, — r/comfyui
  4. DeepL is now available in Microsoft Copilot, ChatGPT and Claude via MCP — DeepL Blog
  5. Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test? — How I AI
  6. Claude is BACK! — r/ClaudeAI
  7. AI slowdown? OpenAI and Anthropic launch cheaper AI models days after urging caution on frontier AI — Mint AI
  8. Anthropic and OpenAI roll out cheaper models in first release since call for slowdown — CNBC Technology

Get the daily brief of stories like this at 6:30 every morning →