A key strength of Proximal Policy Optimization (PPO) is its learned critic, which uses historical trajectories collected during reinforcement learning to est…
机构:伯克利
来源:arXiv 2609.36802 | AI4Papers 论文推荐平台