Reinforcement learning (RL) improves reasoning, but its performance depends on the policy from which training begins. We study on-policy distillation (OPD) a…
机构:JD.COM
来源:arXiv 2609.28145 | AI4Papers 论文推荐平台