Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder …
来源:arXiv 2610.03389 | AI4Papers 论文推荐平台