Vision-language models (VLMs) are increasingly trained to generate structured outputs like points and bounding boxes that downstream interfaces, agents, and …
机构:UC Riverside
来源:arXiv 2610.05637 | AI4Papers 论文推荐平台