Long-video agents can actively gather question-relevant evidence, but they typically leave a central decision implicit: when has the agent seen enough to ans…
机构:北大
来源:arXiv 2610.05400 | AI4Papers 论文推荐平台