Serving long documents to a Large Language Model (LLM) repeatedly is expensive: computations grow with context length, and the memory footprint of the key-va…
机构:Apple
来源:arXiv 2609.35621 | AI4Papers 论文推荐平台