Advances in language modeling have been driven by scaling pretraining on ever more data. Yet, the training data is still largely curated on the model's behal…
机构:Stanford
来源:arXiv 2609.30063 | AI4Papers 论文推荐平台