Sparse autoencoders (SAEs) are an important tool for mechanistic interpretability, but interpreting their many features remains challenging. Existing methods…
机构:清华
来源:arXiv 2609.35367 | AI4Papers 论文推荐平台