Increasing the batch size during training is a common practice in large language model (LLM) pretraining, yet the theoretical justification behind its succes…
机构:Purdue University
来源:arXiv 2610.02355 | AI4Papers 论文推荐平台