A lot of prior work addressed key-value (KV) cache selection and compression by sparse attention to enable long-context inference for transformer language mo…
机构:Amazon Web Services
来源:arXiv 2608.19920 | AI4Papers 论文推荐平台