[P] τ²-bench's outcome-only reward taught our GRPO support agent to hand off to humans whenever unsure
This is a course project, but the failure mode seemed worth sharing. Setup: Qwen3-4B-Instruct-2507, SFT warm start, then multi-turn GRPO on τ²-bench airline + retail. τ²-bench has almost no hand-off tasks, so we derived 461 of them (272 should-transfer, 189 hard negatives) from the upstream tasks,…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-10-10 21:29 · r/learnmachinelearning
[P] τ²-bench's outcome-only reward taught our GRPO support agent to hand off to humans whenever unsure