Off-policy hierarchical reinforcement learning must estimate the values of high-level decisions while the low-level policy changes. HIRO adapts replay data t…
机构:中科院
来源:arXiv 2610.05446 | AI4Papers 论文推荐平台