Reinforcement learning with verifiable rewards (RLVR) has improved the reasoning capabilities of large language models (LLMs), yet their predictions remain s…
机构:浙大
来源:arXiv 2609.40360 | AI4Papers 论文推荐平台