If you repair every malformed sample, what policy are you evaluating?
Silently asking a model to try again until its output parses changes the experiment. The system's success rate now includes a retry policy that the first sampled answer didn't earn. The Guidance-TTT example in Reef Infra makes a different choice. A trainable model produces structured guidance, and…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-10-11 17:32 · r/reinforcementlearning
If you repair every malformed sample, what policy are you evaluating?