Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual token…
机构:南大
来源:arXiv 2609.37581 | AI4Papers 论文推荐平台