Reinforcement learning with verifiable rewards (RLVR) supervises mathematical reasoning through final-answer correctness, but provides little guidance on ind…
机构:上交
来源:arXiv 2609.28385 | AI4Papers 论文推荐平台