Critic-free reinforcement fine-tuning (RFT) for agentic large language models is often done through GRPO-style methods, which compute a group baseline over r…
机构:剑桥
来源:arXiv 2610.03361 | AI4Papers 论文推荐平台