Sparse attention reduces the cost of long-context attention, but existing kernels typically process queries independently, repeatedly loading and dequantizin…
机构:华为
来源:arXiv 2610.11134 | AI4Papers 论文推荐平台