On-policy distillation (OPD) has emerged as an effective framework for post-training language models by pairing student-generated trajectories with dense tok…
机构:香港科技大学(广州)
来源:arXiv 2608.19408 | AI4Papers 论文推荐平台