We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-re…
来源:arXiv 2609.35505 | AI4Papers 论文推荐平台