While most attention logits can be computed in low precision without degrading numerical stability, current attention kernels fail to exploit this phenomenon…
机构:维也纳大学
来源:arXiv 2609.39361 | AI4Papers 论文推荐平台