Long-output reasoning shifts the KV-cache bottleneck from the fixed prompt to the generated trace. Existing reasoning-cache eviction methods largely treat ca…
机构:LinkedIn
来源:arXiv 2610.03007 | AI4Papers 论文推荐平台