Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action, yet fai…
机构:清华
来源:arXiv 2609.39971 | AI4Papers 论文推荐平台