Reward post-training of flow models either reweights the model's own samples under a KL penalty or a frozen reference, often for thousands of updates, or bac…
机构:Google
来源:arXiv 2610.05954 | AI4Papers 论文推荐平台