Existing in-context reinforcement learning methods mainly pretrain Transformers with supervised behavior-prediction objectives. This enables task inference f…
机构:上交
来源:arXiv 2609.30391 | AI4Papers 论文推荐平台