Mixture-of-Experts (MoE) architectures have become essential for scaling large language models, with recent work demonstrating the benefits of fine-grained e…
机构:新加坡国立大学
来源:arXiv 2610.11317 | AI4Papers 论文推荐平台