Vision-language models face a fundamental scaling bottleneck: the number of visual tokens grows with both temporal duration and spatial resolution, making lo…
机构:百度
来源:arXiv 2609.39924 | AI4Papers 论文推荐平台