Foundation models supply reinforcement learning (RL) with priors that mitigate its longstanding weaknesses in sample efficiency and transfer, but their token…
机构:南大
来源:arXiv 2610.11692 | AI4Papers 论文推荐平台