Post-training quantization (PTQ) has become a widely adopted technique for reducing the memory footprint and inference cost of large language models (LLMs). …
机构:上交
来源:arXiv 2609.39801 | AI4Papers 论文推荐平台