Distillation has become a core primitive of large language model training, but its properties are not yet well understood. We take an entropic perspective, s…
机构:Stanford
来源:arXiv 2610.03529 | AI4Papers 论文推荐平台