One repo, two loops: the one that trains nothing and the one that is RL
Half the "self-improving agent" posts get the reply "that's not RL, that's prompt tuning", and half the time the reply is right. reef has both loops in one codebase and I run the harness side, so here's where the line sits. The harness loop needs a model endpoint and no GPU. Its recipes (Reefine, S…
Read the full story at r/reinforcementlearning ↗
Timeline · 1 report
- 2026-10-02 04:26 · r/reinforcementlearning
One repo, two loops: the one that trains nothing and the one that is RL