Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning
This story is from 2026-09-02. It is preserved in the archive; the latest stories are on the live feed.
arXiv:2609.00213v1 Announce Type: new Abstract: Reinforcement learning fine-tuning of large language models increasingly adopts multiple reward dimensions, including verifiable rules, task-specific evaluators, and learned reward models, to provide richer supervision across diverse capabilities. The…
Read the full story at arXiv cs.CL ↗
Timeline · 1 report
- 2026-09-02 04:00 · arXiv cs.CL
Uncovering and Mitigating Aggregation-Induced Reward Hacking in Multi-Reward Reinforcement Learning