Many LLM inference problems, including model routing, prefix-cache management, prompt trimming, and test-time search, can be viewed as optimization over a tr…
机构:Amazon
来源:arXiv 2609.30017 | AI4Papers 论文推荐平台