Tool-using language-model agents are vulnerable to indirect prompt injection because they must act on untrusted external content. Existing training-time defe…
机构:巴黎综合理工
来源:arXiv 2610.06401 | AI4Papers 论文推荐平台