FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience
This story is from 2026-09-04. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.03241v1 Announce Type: new Abstract: A reasoning model can improve from its own on-policy experience, but this inner loop is fragile: terminal verifiers provide reliable yet sparse supervision, while dense same-model guidance can reinforce false confidence or overconcentrate learning on…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-04 04:00 · arXiv cs.LG
FlowBalance: Verifier-Grounded Self-Improvement from On-Policy Reasoning Experience