Current vision-language models (VLMs) excel at visual content understanding and text-based reasoning, yet their structure limits the advancement of incorpora…
机构:麻省大学阿默斯特分校
来源:arXiv 2609.30130 | AI4Papers 论文推荐平台