Adding memory to search instead of sampling in reward maximization tasks [R]
I am one of the authors of FLEET - an algorithm that enhances Best-of-N generation by attributing external rewards to particular tokens and then uses MCTS to adjust logits during the next run. I find it rather funny that most of the tasks where repetitive sampling is widely used are based on reward…
Read the full story at r/MachineLearning ↗
Timeline · 1 report
- 2026-10-02 12:04 · r/MachineLearning
Adding memory to search instead of sampling in reward maximization tasks [R]