On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors …
机构:清华
来源:arXiv 2609.37522 | AI4Papers 论文推荐平台