Despite their impressive performance on a wide range of video understanding tasks, current Vision Language Models (VLMs) are predominantly designed for offli…
机构:南洋理工
来源:arXiv 2610.05940 | AI4Papers 论文推荐平台