Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction…
机构:康奈尔大学
来源:arXiv 2609.40221 | AI4Papers 论文推荐平台