Agentic reinforcement learning (RL) has become a critical stage in the post-training of large language models. Existing critic-free, group-relative methods e…
机构:厦门大学
来源:arXiv 2608.19842 | AI4Papers 论文推荐平台