Most KV-cache compression methods classify attention heads once, either offline or during prefill, and keep this classification fixed throughout generation. …
机构:UCSD
来源:arXiv 2610.02713 | AI4Papers 论文推荐平台