On-policy distillation (OPD) trains students using teacher feedback on their own sampled responses, yet how prompt choice shapes transfer across teacher-stud…
机构:腾讯
来源:arXiv 2609.37377 | AI4Papers 论文推荐平台