Data selection is already a central bottleneck in large-language-model training, where web-scale corpora are noisy and token budgets are finite. In continual…
机构:UIUC
来源:arXiv 2610.02593 | AI4Papers 论文推荐平台