We answer this question with Vector-Symbolic Policy Gradient (VSPG), a discrete-action actor that represents each action by a unit-norm hypervector and score…
机构:UCI
来源:arXiv 2608.18404 | AI4Papers 论文推荐平台