Long-horizon LLM agents are typically trained with sparse outcome rewards, making trajectory-level objectives too coarse to distinguish the contribution of i…
机构:小米
来源:arXiv 2610.03223 | AI4Papers 论文推荐平台