On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full…
机构:MIT
来源:arXiv 2609.39275 | AI4Papers 论文推荐平台