As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts…
机构:Apple
来源:arXiv 2609.37988 | AI4Papers 论文推荐平台