Group Relative Policy Optimization (GRPO) is widely used to train reasoning language models, where it computes advantages by centering and normalizing reward…
机构:港科大
来源:arXiv 2609.36820 | AI4Papers 论文推荐平台