xgbj/sparse-mask-attention

高性能短序列稀疏Mask Attention CUDA算子,针对<1K序列+75%稀疏度优化

80Stars on the repository
1Mods indexed here, across every type
5mo agoLast push, which is what freshness is scored on
noneNo LICENSE: all rights reserved, so bodies are not copied