Standard language model RL algorithms credit every token of a long rollout with the same advantage determined by the terminal reward. Actor-critic methods ca…
机构:Stanford
来源:arXiv 2609.39247 | AI4Papers 论文推荐平台