Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully di…
机构:爱丁堡大学
来源:arXiv 2609.31571 | AI4Papers 论文推荐平台