Principled exploration in reinforcement learning requires an agent to quantify its epistemic uncertainty and act to resolve it. Uncertainty over the value fu…
机构:苏黎世联邦理工
来源:arXiv 2609.35525 | AI4Papers 论文推荐平台