How many prompts does on-policy distillation (OPD) need, and how does the answer depend on the student policies that generate its training responses? We stud…
机构:腾讯
来源:arXiv 2609.25048 | AI4Papers 论文推荐平台