Reasoning models often continue generating after their answers have settled. Settle learns when to stop from answer stability in completed traces. It trains …
机构:牛津
来源:arXiv 2609.38997 | AI4Papers 论文推荐平台