Reliable AI safeguards require both control mechanisms that reduce unsafe behavior and monitoring mechanisms that detect safety risks during model interactio…
机构:北大
来源:arXiv 2609.34771 | AI4Papers 论文推荐平台