Reinforcement learning for long-horizon agents relies on purely retrospective training signals: credit is assigned only after observing environmental consequ…
机构:Salesforce
来源:arXiv 2610.02740 | AI4Papers 论文推荐平台