Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images …
机构:中科院
来源:arXiv 2610.01908 | AI4Papers 论文推荐平台