Safe reinforcement learning commonly places safety and task performance in the same policy objective, where they can introduce competing updates. Safety filt…
机构:MIT
来源:arXiv 2610.12432 | AI4Papers 论文推荐平台