Extreme low-bit LLM compression is most challenging when weights, activations, and KV caches are quantized together: their distributions differ, and quantiza…
机构:Meta
来源:arXiv 2610.09202 | AI4Papers 论文推荐平台