Recent agent benchmarks increasingly ground evaluation in executable environments, from code repair to web navigation, app APIs, and function calling. Yet co…
机构:Microsoft
来源:arXiv 2608.19741 | AI4Papers 论文推荐平台