Reinforcement learning (RL) is the central training paradigm for advancing large foundation models towards self-improvement. This report introduces the MiMo-…
机构:小米
来源:arXiv 2610.11959 | AI4Papers 论文推荐平台