WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps
arXiv:2609.27033v1 Announce Type: new Abstract: Reward fine-tuning aims to update a pre-trained flow-based generative model to improve the downstream reward of its generated samples. Existing methods typically formulate this problem as sampling from a reward-tilted distribution, the solution to a K…
Read the full story at arXiv cs.LG ↗
Timeline · 1 report
- 2026-09-24 04:00 · arXiv cs.LG
WTF?! Simulation-Free Reinforcement Learning with Wasserstein-Tilted Flow Maps