AINewsnow

[P] τ²-bench's outcome-only reward taught our GRPO support agent to hand off to humans whenever unsure

This is a course project, but the failure mode seemed worth sharing. Setup: Qwen3-4B-Instruct-2507, SFT warm start, then multi-turn GRPO on τ²-bench airline + retail. τ²-bench has almost no hand-off tasks, so we derived 461 of them (272 should-transfer, 189 hard negatives) from the upstream tasks,…

Read the full story at r/learnmachinelearning ↗

Timeline · 1 report

  1. 2026-10-10 21:29 · r/learnmachinelearning
    [P] τ²-bench's outcome-only reward taught our GRPO support agent to hand off to humans whenever unsure

More stories

  1. An Anthropic AI model sent a false homicide tip to Philadelphia police — TechCrunch AI
  2. Anthropic bans users from ‘needless abusive or cruel behavior’ towards Claude — The Guardian AI
  3. Philadelphia police receive false homicide tip from Anthropic AI model — The Hill Technology
  4. Nvidia in talks to acquire US ‘open’ model start-up Reflection AI — Financial Times AI
  5. Impactful scheduling for GPU clusters — Allen Institute for AI (Ai2)
  6. Welcome to Gemini at Work 2026: Introducing the Gemini agent — Google Cloud AI Blog
  7. Qwen Image 2.1 Turbo Released -- Hugging Face — r/StableDiffusion
  8. Google Cloud introduces Gemini agent to change enterprise work — SiliconANGLE AI

Get the daily brief of stories like this at 6:30 every morning →