Long-context evaluations often test whether a model can recover distant evidence, but recoverability does not guarantee behavioral influence. We test the pre…
机构:剑桥
来源:arXiv 2609.27590 | AI4Papers 论文推荐平台