On-policy self-distillation (OPSD) trains a language model to match a copy of itself conditioned on privileged context. Existing work varies what privileged …
机构:Stanford
来源:arXiv 2610.07842 | AI4Papers 论文推荐平台