As large language model (LLM) agents become increasingly autonomous, they may pursue task performance through deception, raising concerns about their reliabl…
机构:西安交大
来源:arXiv 2610.07967 | AI4Papers 论文推荐平台