Tropical Reinforcement Learning
arXiv:2610.02478v1 Announce Type: new Abstract: Reinforcement learning for large language models typically maximizes expected return, adding up the probabilities of all successful trajectories. However, the classical sum formulation can only report how often the model policy succeeds, not which sol…
Read the full story at arXiv cs.AI ↗
Timeline · 1 report
- 2026-10-05 04:00 · arXiv cs.AI
Tropical Reinforcement Learning