Large language models (LLMs) can sometimes report perturbations to their internal activations---even when the input provides no evidence that an intervention…
机构:清华
来源:arXiv 2609.35108 | AI4Papers 论文推荐平台