While recent vision-language models (VLMs) have shown outstanding performance across diverse applications, they tend to under-use visual information and over…
机构:Amazon
来源:arXiv 2610.08401 | AI4Papers 论文推荐平台