The large KV-cache size of modern LLMs creates a barrier to efficient deployment. Recent work has explored replacing attention layers' RoPE positional embedd…
机构:IBM
来源:arXiv 2610.09051 | AI4Papers 论文推荐平台