False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Self-evolving search agents can produce misleading training signals when a question proposer and solver learn from the same pseudo-labels. Their agreement may increase because they share errors, even while correctness against the source material stagnates. This paper calls that failure mode “co-che…
Read the full story at r/learnmachinelearning ↗
Timeline · 1 report
- 2026-10-08 19:29 · r/learnmachinelearning
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents