During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe lat…
机构:北大
来源:arXiv 2610.02835 | AI4Papers 论文推荐平台