Reinforcement learning with verifiable rewards provides a sparse post-training signal: a single binary outcome evaluates the entire rollout, and every token …
机构:北大
来源:arXiv 2609.34849 | AI4Papers 论文推荐平台