Reinforcement learning (RL) methods such as GRPO substantially improve large language model reasoning but often suffer from policy entropy collapse: the loss…
机构:哈工大
来源:arXiv 2610.11854 | AI4Papers 论文推荐平台