Data Science Wire

EvoRS: On-Policy Self-Evolution of Reward Systems for Open-Ended Reinforcement Learning

arXiv cs.AI1d4 min read

arXiv:2609.12459v1 Announce Type: new Abstract: Open-ended reinforcement learning often relies on rubric-based rewards for tasks without directly verifiable answers. Yet the policy and reward system form a dynamic feedback loop: as the policy optimizes the current reward, an initially useful reward system may become unreliable due to reward hacking or reduced response discriminability. The reward system should therefore evolve rather than remain fixed during training. Existing dynamic-rubric methods adapt evaluation criteria, but reward failures can also arise from scoring mechanisms or signal

Read the full story at arXiv cs.AI

More in Machine Learning