On-policy distillation (OPD) trains a student model by having it generate trajectories, then matching its next-token predictions with an external teacher's n…
机构:Meta
来源:arXiv 2609.30652 | AI4Papers 论文推荐平台