Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/huggingface/kernels/cpu-kernelsnpx skills add huggingface/kernels --skill cpu-kernelsgit clone --depth 1 https://github.com/huggingface/kernelsWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00089 | $0.06865 |
| Opus 5 | $0.00044 | $0.03432 |
| Sonnet 5 | $0.00018 | $0.01373 |
| Haiku 4.5 | $0.00009 | $0.00686 |
Grade A, and why
cpu-kernels scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 462 lines — stays where its author put it; the contents beside it link to each section on GitHub.
CPU C++ Kernels for x86 Processors
This skill provides patterns and guidance for developing optimized C++ kernels targeting x86 CPUs (Intel Xeon and compatible processors) with AVX2 and AVX512 intrinsics. Kernels are compiled via kernel-builder and distributed through the Hugging Face kernels ecosystem.
Who runs these commands? You, the agent — not a human. This is an autonomous loop: you write/edit the C++ kernel, build it, then run the scripts below as tools (via Bash) to check correctness, benchmark, and profile. You read each result, record it with
trial_manager.py, decide the next change from the Phase 2 decision tree, and repeat until you hitearly_stop_speedupor run allmax_trials.
Key Concepts (read before the Quick Start)
The commands use a few names that mean different things. They are not interchangeable:
| Name (example) | What it is | Used by |
|---|---|---|
baseline.py |
The PyTorch reference implementation you optimize against. It is the ground truth for correctness and the speed reference for speedup. It must define get_inputs() and either get_reference_output() or a Model class (plus optional get_init_inputs()). You write this file (or it is given) before starting. |
every script |
my_rmsnorm |
A trial-tree label — an arbitrary name you pick for this optimization task. trial_manager.py stores all attempts under trials/my_rmsnorm/. It is only a tracking ID. |
trial_manager.py only |
my_kernel |
The installed Python package name — the build artifact produced by kernel-builder build + pip install. This is the importable module that contains your compiled kernel. |
--kernel-package |
my_kernel.rms_norm |
An <package>.<function> path — the actual callable inside the installed package. Passed to --op to tell the benchmark/profiler which function to run. |
--op |
⚠️
--opmeans two different things depending on the script. Inanalyze_op.py,--opis a plain operation name (e.g."rms_norm") used to look up compute/memory characteristics. Inbenchmark_cpu.pyandcpu_profiler.py,--opis apackage.functionpath (e.g.my_kernel.rms_norm) used to import and call your kernel. Same flag, different meaning — read each command below carefully.
What ships with it
22 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- manifest.txt 685 B
- references/brgemm_patterns.yaml 14 KB
- references/build_system.yaml 7.8 KB
- references/correctness.yaml 5.7 KB
- references/dtype_optimizations.yaml 3.4 KB
- references/huggingface-kernels-integration.md 4.4 KB
- references/implementation_reference.md 12 KB
- references/memory_patterns.yaml 3.4 KB
- references/optimization_levels.yaml 5.8 KB
- references/optimization_strategies.md 4.2 KB
- references/quantized_gemm_patterns.yaml 18 KB
- references/runtime_dispatch.yaml 8.6 KB
- references/simd_optimization_patterns.yaml 6.9 KB
- references/threading_patterns.yaml 3.0 KB
- references/workflow_details.md 7.9 KB
- scripts/analyze_op.py 11 KB runs code
- scripts/benchmark_cpu.py 8.8 KB runs code
- scripts/config.py 635 B runs code
- scripts/config.yaml 641 B
- scripts/cpu_profiler.py 10 KB runs code
- scripts/trial_manager.py 14 KB runs code
- scripts/validate_cpu_kernel.py 12 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 462 lines · 89 tokens per session scan A b7686f78f8b9
cpu-kernels is a skill published in the GitHub repository huggingface/kernels (729 stars, last pushed 4d ago), licensed Apache-2.0. It adds 89 tokens to every session and 6,865 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
babysit-pr
Babysit a GitHub pull request after creation by continuously polling review comments, CI checks/workflow runs, and mergeability state until the PR is merged/closed or user help is required. Diagnose failures, retry likely flaky failures up to 3 times, auto-fix/push branch-related issues when appropriate, and keep…
imagegen
Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output…
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…
next-cache-components-optimizer
Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…