Vision-language models (VLMs) may accept false visual premises, answering questions about a target object's color, count, location, or state even when it is …
机构:UCLA
来源:arXiv 2609.38111 | AI4Papers 论文推荐平台