Large language model (LLM) inference is constrained by the quadratic scaling of self-attention and the linear scaling of the KV cache, increasing latency, en…
机构:Notre Dame
来源:arXiv 2609.25537 | AI4Papers 论文推荐平台