Large language models rely on KV caching to reduce repeated prefill computation in long context and interactive applications. As lightweight adapters evolve,…
机构:南大
来源:arXiv 2609.30884 | AI4Papers 论文推荐平台