Language-model agents increasingly face long-horizon tasks with evolving state, interdependent decisions, and delayed outcomes. Scaling their training requir…
机构:阿里
来源:arXiv 2609.27321 | AI4Papers 论文推荐平台