Multi-reward policy optimization requires a joint update that reflects both the learning signals and the intended relationships among objectives. We introduc…
机构:复旦
来源:arXiv 2609.34985 | AI4Papers 论文推荐平台