General-purpose vision-language models (VLMs) bring broad knowledge and spatial reasoning to robot manipulation, yet existing systems either use them indirec…
机构:HKUST(GZ)
来源:arXiv 2609.29964 | AI4Papers 论文推荐平台