Mixture-of-Experts (MoE) has been widely adopted in recent large language model (LLM) architectures. However, scaling up MoE in LLM training introduces syste…
机构:Cornell
来源:arXiv 2610.05744 | AI4Papers 论文推荐平台