AINewsnow

Claude Code's /goal judge said "done" in 17 of 17 runs. 8 of them were broken.

This story is from 2026-10-04. It is preserved in the archive; the latest stories are on the live feed.

I like /goal in Claude Code. You say what "done" looks like and Claude keeps working until it gets there. What bothered me is who decides it got there. After every turn a small model (Haiku) reads the conversation and answers "met", "not yet" or "impossible". I tested what that judge can see. It ne…

Read the full story at DEV Community — AI ↗

Timeline · 1 report

  1. 2026-10-04 12:22 · DEV Community — AI
    Claude Code's /goal judge said "done" in 17 of 17 runs. 8 of them were broken.

More stories

  1. Building ambient agents with Amazon Bedrock AgentCore: From event-driven signals to human-in-the-loop workflows — AWS Machine Learning Blog
  2. Kimi K3: A Claude clone or something else? — CoreWeave Blog
  3. Implementing Multi-Environment Access for Claude Platform on AWS — AWS Machine Learning Blog
  4. The Museum of Lost Things | Short Film by Claude (Minimax H3) NO user input. — r/ClaudeAI
  5. If you have subscription of both, this will let your Claude Code and Codex collaborate much better. — r/ChatGPT
  6. An AI couldn’t beat humans at StarCraft, so it decided to cheat — The Verge AI
  7. DeepSeek Harness v0.2 Brings Official Desktop Apps to Its Open-Source Agent Harness — MarkTechPost
  8. How do I stop feeling left behind with Gemini Pro? (Trying to replicate Claude Code / homelab setups) — r/Bard

Get the daily brief of stories like this at 6:30 every morning →