Visual grounding localizes an object described by language with a bounding box. Most multimodal grounding models compress target identification, spatial reas…
机构:西安交大
来源:arXiv 2610.01710 | AI4Papers 论文推荐平台