An emerging hypothesis about reinforcement learning (RL) post-training of large language models (LLMs) is that it merely sharpens existing behaviors of a bas…
机构:Meta
来源:arXiv 2610.01509 | AI4Papers 论文推荐平台