Online reinforcement learning fine-tuning of pretrained flow-matching vision-language-action (VLA) policies promises robots that keep learning after deployme…
机构:卡尔斯鲁厄理工学院
来源:arXiv 2609.28838 | AI4Papers 论文推荐平台