Majority voting over sampled completions is the workhorse of test-time scaling, and reinforcement learning with verifiable rewards (RLVR) is the workhorse fo…
机构:普林斯顿大学
来源:arXiv 2610.00991 | AI4Papers 论文推荐平台