As autonomous coding agents get increasingly deployed, the risk that accidental or adversarially injected misspecifications in tasks lead to dangerous agent …
机构:DeepMind
来源:arXiv 2610.09159 | AI4Papers 论文推荐平台