KV-cache quantization and linear attention are two representative approaches to tackling the storage and computational costs of Transformers. KV-cache quanti…
机构:港中文
来源:arXiv 2610.11214 | AI4Papers 论文推荐平台