Multimodal large language models have become the dominant paradigm for visual understanding, but incur substantial costs by encoding inputs into dense, fixed…
机构:阿里
来源:arXiv 2610.07987 | AI4Papers 论文推荐平台