Explicit intermediate reasoning gives large language models (LLMs) a stronger problem-solving mode. We study learning from this think-mode advantage via on-p…
机构:字节
来源:arXiv 2609.37044 | AI4Papers 论文推荐平台