Reinforcement learning (RL) for post-training large language models (LLMs) incurs substantial computation and memory overhead during rollout generation, whic…
机构:阿里
来源:arXiv 2610.07767 | AI4Papers 论文推荐平台