Multi-turn rollout dominates the cost of agentic reinforcement learning (RL). Asynchronous execution and elastic GPU resources can accelerate this stage, but…
机构:USTC
来源:arXiv 2609.35158 | AI4Papers 论文推荐平台