Self-attention gives LLMs fine-grained, query-dependent access to context, but dense token interactions incur quadratic prefill cost and a key--value cache g…
机构:阿里
来源:arXiv 2609.39661 | AI4Papers 论文推荐平台