Long-horizon tasks with sparse rewards pose an exploration bottleneck for goal-conditioned reinforcement learning: a policy started from the initial state ra…
机构:ETH Zurich
来源:arXiv 2610.03395 | AI4Papers 论文推荐平台