AINewsnow

I read the GPT-6 Astra system card and I think we all misunderstand what "monitorability" means now

I'm not a safety researcher, just build stuff with LLM's. Read some of the actual 30 page card, not the summary, and one section keeps bugging me. They measured whether Astra can sandbag. Told it "underperform on this evaluation," then checked if their monitors could catch it. Model dropped from 84…

Read the full story at r/artificial ↗

Timeline · 1 report

  1. 2026-09-27 23:58 · r/artificial
    I read the GPT-6 Astra system card and I think we all misunderstand what "monitorability" means now

More stories

  1. Unsecured OpenAI agents posted 53 user images on the internet without the lab's knowledge — TechCrunch AI
  2. new update? — r/GeminiAI
  3. Opus 5.5 vs. GPT-6 Sol: which model won my blind taste test? — How I AI
  4. Opus 5.5 — r/ClaudeAI
  5. Proaction boosts sales 60% and saves 75+ hours with Codex — OpenAI News
  6. Need Help & Advice !! — r/AI_Agents
  7. AI LEARNING QUESTION — r/learnmachinelearning
  8. Optimizing my AI subscriptions: Claude Pro (Opus) vs. ChatGPT Plus vs. Perplexity Pro? — r/AI_Agents

Get the daily brief of stories like this at 6:30 every morning →