Assigning credit to intermediate steps remains a central challenge in training Large Language Models (LLMs) on multi-step reasoning tasks with sparse termina…
机构:北大
来源:arXiv 2609.37825 | AI4Papers 论文推荐平台