Language model (LM) alignment broadly aims to perturb a given LM $Q$ into an aligned LM $q$ such that i) the outputs produced by $q$ and $Q$ are 'close' in p…
机构:Georgia Tech
来源:arXiv 2610.01828 | AI4Papers 论文推荐平台