Large language models perform strongly on competition mathematics, but their research-level reasoning remains difficult to evaluate systematically. Theoretic…
机构:Amazon
来源:arXiv 2609.35606 | AI4Papers 论文推荐平台