On-policy reinforcement learning has become a central paradigm for improving the reasoning abilities of large language models. However, its effectiveness is …
机构:NYU
来源:arXiv 2610.07342 | AI4Papers 论文推荐平台