Reinforcement learning (RL) has become a cornerstone for improving the reasoning capabilities of large language models (LLMs), but the need for on-policy dat…
机构:MIT
来源:arXiv 2609.36816 | AI4Papers 论文推荐平台