xgbj/sparse-mask-attention

A high-performance CUDA operator for sparse Mask Attention on short sequences, optimized for sequences under 1K with 75% sparsity

1 file for Claude Code: sparse-mask-attention CLAUDE.md — 653 tokens loaded in every session.

80Stars on the repository
1Files it configures its agents with
653Tokens loaded in every session
1Agent configured

Instructions

These files are xgbj/sparse-mask-attention's own configuration — they tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.