Scaling the horizon of agentic LLMs is bottlenecked by the need to fit ever longer context traces in GPU memory. Context compaction has been the most popular…
机构:Mila
来源:arXiv 2609.35750 | AI4Papers 论文推荐平台