Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield polici…
机构:UIUC
来源:arXiv 2610.01133 | AI4Papers 论文推荐平台