GPU skills

96 tagged GPU, measured the same way as everything else here.

Browse within: deep-learning 26cuda 22autograd 16infrastructure 16neural-network 16numpy 16tensor 16NVIDIA 12cloud-functions 12Orchestration 11inference 11kubernetes 11serverless-gpu 10zymtrace 10

fix-ci

25

CliMA/CloudMicrophysics.jl

Skill Claude CodeCodex

Fix CI failures and performance regressions for a Julia package PR by iterating - triage the latest CI results, fix one root cause, verify locally, push. Use when a PR's GitHub Actions or Buildkite CI is failing, or when CI jobs run slower than they do on the main branch. Verifies CPU and GPU compilation locally…

not rated 53 8d ago A 84 tokens copy · 100% Apache-2.0

swm-gpu-workflow

26

swm-gpu/swm

Skill Claude CodeCodex

Use when the user wants to search cloud GPUs across providers, provision a GPU pod, install AI frameworks (ComfyUI, SwarmUI, vLLM, Ollama, Open WebUI, Axolotl, H2O LLM Studio), pull models from HuggingFace / Civitai / URLs to a unified store, sync workspaces to S3, manage idle-pod lifecycle, or track GPU spend — all…

not rated 24 11d ago A 116 tokens original Apache-2.0

sweet-index

27

mrsladoje/sweet-search

Skill Claude CodeCodex

Use when (re)indexing a Sweet Search project. Runs the full-profile indexer with GPU model prewarming (CoreML cascade on M3+, candle Metal on M1/M2, ORT CPU elsewhere), kills ORT CPU models during indexing to avoid memory contention, and rewarms them for query readiness on completion. Incremental runs under 20 files…

not rated 21 6d ago A 83 tokens original Apache-2.0

fix-ci

28

CliMA/SurfaceFluxes.jl

Skill Claude CodeCodex

Fix CI failures and performance regressions for a Julia package PR by iterating - triage the latest CI results, fix one root cause, verify locally, push. Use when a PR's GitHub Actions or Buildkite CI is failing, or when CI jobs run slower than they do on the main branch. Verifies CPU and GPU compilation locally…

not rated 20 7d ago A 84 tokens copy · 100% Apache-2.0

howdeploy/deploychan_mcp

Skill Claude CodeCodex

Use when the user wants to run ComfyUI image/video generation on a rented remote GPU instead of paid API subscriptions: rent a GPU box, bootstrap the stack, connect the agent over SSH, drive workflows via the ComfyUI API, and deliver results to Telegram.

not rated 11 5d ago A 68 tokens original MIT

local-search

30

KempnerInstitute/hpc-agentic-recipes

Skill Claude CodeCodex

Search the web, arxiv, Crossref, PubMed, OpenAlex, or Wikipedia, and fetch page text, using common/tools/search.sh. Use whenever the user asks to search online, look something up, or find papers while served by a local model, and whenever the built-in WebSearch tool fails or is silently dropped by the endpoint.

not rated 11 +1 7d ago A 73 tokens original MIT

faster-whisper

31

ThePlasmak/faster-whisper

Skill Claude CodeCodex

Local speech-to-text using faster-whisper. 4-6x faster than OpenAI Whisper with identical accuracy; GPU acceleration enables 20x realtime transcription. SRT/VTT/TTML/CSV subtitles, speaker diarization, URL/YouTube input, batch processing with ETA, transcript search, chapter detection, per-file language map.

not rated 11 +1 6mo ago A 74 tokens original MIT

troycheng/cuda-kernel-optimizer

Skill Claude CodeCodex

Use when optimizing, tuning, diagnosing, or profiling CUDA, CUTLASS, Triton, PyTorch, vLLM, TensorRT-LLM, or another GPU workload; when assessing an NCU, Nsys, or PyTorch Profiler report; or when the test workload, correctness checks, measurement path, or target environment is incomplete.

not rated 7 6d ago A 77 tokens original MIT

vastai-gpu

33

Yusuke710/vastai-skill

Skill Claude CodeCodex

Rent a GPU on vast.ai, run experiments on it over SSH/rsync, pull artifacts back, and destroy it. Use when a task needs a GPU (training, CUDA, large-model inference).

not rated 3 1mo ago A 46 tokens

mlx-metal-kernels

34

ssmall256/mlx-metal-kernels-skill

Skill Claude CodeCodex

Provides guidance for writing custom Metal compute kernels using MLX's mx.fast.metalkernel() API for Apple Silicon GPUs (M1, M2, M3, M4). Covers kernel patterns for RMSNorm, LayerNorm, softmax, attention variants (causal, sliding window, GQA), and reductions. Includes profiling, debugging, simdgroupmatrix MMA (M3+)…

not rated 2 6mo ago A 91 tokens original MIT

autodl

35

wuzihuang/AUTODL-PLUGIN

Skill Claude CodeCodex

Use AutoDL safely through its official developer APIs and documentation. Trigger for AutoDL balance, Container Instance Pro, elastic deployment, images, GPU stock, NFS, duration packages, WeChat notifications, container metrics, billing, storage, data retention, SSH/JupyterLab/VSCode, CUDA environments…

not rated 2 14d ago A 73 tokens original MIT

project-bourne

36

KozakHou/project-bourne

Skill Claude CodeCodex

Use Project Bourne to plan, execute, reproduce, inspect, or trace scientific and engineering workloads when durable provenance matters. Trigger for simulations, numerical solvers, ML or training, GPU and HPC runs on local compute, Slurm, PBS, or LSF, reproducible experiments, failed attempts, experiment comparisons…

not rated 0 5d ago A 91 tokens original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: