Deploying Mixture-of-Experts (MoE) models relies heavily on Expert Parallelism, which generates intense inter-GPU communication. Consequently, state-of-the-a…
机构:清华
来源:arXiv 2610.09424 | AI4Papers 论文推荐平台