Video temporal grounding aims to localize events in videos from natural-language queries. For agents built around frozen video-language models, the harness d…
机构:清华
来源:arXiv 2610.01766 | AI4Papers 论文推荐平台