Reinforcement learning with verifiable rewards (RLVR) has substantially improved the mathematical reasoning capabilities of large language models. Recent wor…
机构:Amazon
来源:arXiv 2610.01080 | AI4Papers 论文推荐平台