The key-value (KV) cache becomes a major memory bottleneck in long-context LLM inference, placing substantial pressure on memory capacity and bandwidth. To m…
机构:Stanford
来源:arXiv 2610.03027 | AI4Papers 论文推荐平台