Audio language models understand what is said far better than how it sounds. Closing this gap takes more than data. Detailed acoustic annotation is costly, l…
机构:腾讯
来源:arXiv 2609.27389 | AI4Papers 论文推荐平台