LLM agents often fail mid-task due to invalid tool calls, repeated actions, or poorly grounded reasoning, and learning from these failures is a path to relia…
机构:Stanford
来源:arXiv 2610.02994 | AI4Papers 论文推荐平台