A high-performance CUDA operator for sparse Mask Attention on short sequences, optimized for sequences under 1K with 75% sparsity
1 file for Claude Code: sparse-mask-attention CLAUDE.md — 653 tokens loaded in every session.
CLAUDE.md A 653 tok These files are xgbj/sparse-mask-attention's own configuration — they tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.