Many useful language-model tasks cannot be evaluated by exact outcome verification. Rubric-based reinforcement learning (RL) addresses this issue by scoring …
机构:Meta
来源:arXiv 2610.02781 | AI4Papers 论文推荐平台