Reinforcement learning with verifiable rewards (RLVR) trains reasoning models to produce correct answers, but does not ensure that their stated confidence is…
机构:阿里
来源:arXiv 2609.34857 | AI4Papers 论文推荐平台