Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameter…
机构:Meta
来源:arXiv 2609.40316 | AI4Papers 论文推荐平台