Most preference optimization methods, such as Direct Preference Optimization (DPO), apply preference supervision at the response level, although autoregressi…
机构:NVIDIA
来源:arXiv 2610.01511 | AI4Papers 论文推荐平台