Sparse mixture-of-experts (MoE) large language models scale model capacity by routing each token to a small subset of experts. Their routers are regularized …
机构:NVIDIA
来源:arXiv 2609.37751 | AI4Papers 论文推荐平台