On-policy distillation (OPD) has become an important approach to language model post-training. However, despite its performance gains, OPD can also collapse …
机构:西湖大学
来源:arXiv 2610.03185 | AI4Papers 论文推荐平台