Long-horizon language agents often receive supervision only from terminal task outcomes, leaving little signal for distinguishing productive intermediate beh…
机构:新加坡国立大学
来源:arXiv 2610.09560 | AI4Papers 论文推荐平台