On-policy self-distillation (OPSD) has attracted growing interest as a promising approach to improve the reasoning ability of language models. Without extern…
机构:北大
来源:arXiv 2609.39118 | AI4Papers 论文推荐平台