Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively ex…
机构:NVIDIA
来源:arXiv 2610.01785 | AI4Papers 论文推荐平台