Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have become two main paradigms for post-training reasoning models. RLV…
机构:腾讯
来源:arXiv 2610.11519 | AI4Papers 论文推荐平台