FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models
This story is from 2026-09-10. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.09905v1 Announce Type: new Abstract: Preference alignment for flow and diffusion models now spans online reinforcement learning and offline preference optimization, but the relation between these methods remains unclear. In particular, existing forward-process alignment methods require f…
Read the full story at arXiv stat.ML ↗
Timeline · 1 report
- 2026-09-10 04:00 · arXiv stat.ML
FlowCPO: A Unified Divergence View of Preference Alignment for Flow Models