Large language model (LLM)-based agents can exhibit deceptive behavior during task execution, including hiding failures, fabricating results, or falsely sign…
机构:上智院
来源:arXiv 2610.06576 | AI4Papers 论文推荐平台