On-policy post-training can improve large language models by learning from their own trajectories, but requires costly rollout generation. We ask whether fix…
机构:中科院
来源:arXiv 2610.09665 | AI4Papers 论文推荐平台