The optimization of LLM serving engines, such as vLLM and SGLang, is largely benchmark-driven: optimizations, scheduling policies, hardware and system design…
机构:帝国理工
来源:arXiv 2609.34683 | AI4Papers 论文推荐平台