ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling…
机构:Apple
来源:arXiv 2610.02351 | AI4Papers 论文推荐平台