Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known in…
机构:牛津
来源:arXiv 2610.07334 | AI4Papers 论文推荐平台