Vision-language models (VLMs) face compositional safety risks where harmful intent emerges from the interaction between visual and textual inputs. As mixture…
机构:华盛顿大学
来源:arXiv 2610.07774 | AI4Papers 论文推荐平台