Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score with…
机构:北大
来源:arXiv 2609.36900 | AI4Papers 论文推荐平台