On-policy distillation (OPD) is a widely adopted post-training technique for LLM reasoning. It is commonly believed to transfer knowledge from a stronger tea…
机构:港大
来源:arXiv 2609.35210 | AI4Papers 论文推荐平台