Group-based reinforcement learning such as GRPO trains LLM agents by comparing rollouts sampled for each task, without a learned critic. In long-horizon sett…
机构:北航
来源:arXiv 2609.35082 | AI4Papers 论文推荐平台