Linear attention enables efficient long-context autoregressive decoding by compressing history into recurrent states, but this compression can make selective…
机构:浙大
来源:arXiv 2610.05842 | AI4Papers 论文推荐平台