On-Policy Distillation (OPD) trains a student on its own generations using teacher feedback. With different tokenizers, comparing teacher and student predict…
机构:阿里
来源:arXiv 2610.08448 | AI4Papers 论文推荐平台