Repository-level coding benchmarks have become the standard for evaluating coding agents, yet they inherently suffer from data leakage because they are built…
机构:上交
来源:arXiv 2609.27891 | AI4Papers 论文推荐平台