As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capabil…
机构:Google
来源:arXiv 2610.00972 | AI4Papers 论文推荐平台