Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation. Although sparse attention reduces these costs…
机构:哈工大
来源:arXiv 2609.39068 | AI4Papers 论文推荐平台