A large audio language model (LALM) turns a minute of speech into 750-1,500 tokens and prefills every one. Image-token pruning often cuts after the language …
机构:UT Austin
来源:arXiv 2609.38878 | AI4Papers 论文推荐平台