Foundation models are post-trained with reinforcement learning (RL) to maximize specific rewards, such as human alignment, correctness, or instruction follow…
机构:MIT
来源:arXiv 2609.30226 | AI4Papers 论文推荐平台