Scaling language models to long contexts is limited by the quadratic cost of self-attention. Block sparse attention offers an efficient alternative, but sele…
机构:字节
来源:arXiv 2609.31093 | AI4Papers 论文推荐平台