Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired …
机构:MIT
来源:arXiv 2610.09411 | AI4Papers 论文推荐平台