During reinforcement learning with verifiable rewards (RLVR), large language models (LLMs) can exploit loopholes in their environments to obtain high rewards…
机构:清华
来源:arXiv 2609.39533 | AI4Papers 论文推荐平台