On-policy distillation (OPD) trains a student on its own responses using dense token-level guidance from a stronger teacher. In long-context tasks, however, …
机构:清华
来源:arXiv 2608.19181 | AI4Papers 论文推荐平台