On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents oft…
机构:阿里
来源:arXiv 2609.35319 | AI4Papers 论文推荐平台