Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add OutlineDriven/outline-driven-development --skill cuda-profilinggit clone --depth 1 https://github.com/OutlineDriven/outline-driven-developmentWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/outlinedriven/outline-driven-development/cuda-profiling)<a href="https://agentmods.dev/skills/outlinedriven/outline-driven-development/cuda-profiling"><img src="https://agentmods.dev/badge/skills/outlinedriven/outline-driven-development/cuda-profiling/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/outlinedriven/outline-driven-development/cuda-profiling"><img src="https://agentmods.dev/badge/skills/outlinedriven/outline-driven-development/cuda-profiling.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00042 | $0.01611 |
| Opus 5 | $0.00021 | $0.00805 |
| Sonnet 5 | $0.00008 | $0.00322 |
| Haiku 4.5 | $0.00004 | $0.00161 |
Grade A, and why
cuda-profiling scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 120 lines — stays where its author put it; the contents beside it link to each section on GitHub.
CUDA profiling
Contract
| Field | Bound contract |
|---|---|
| Trigger | A CUDA kernel or pipeline is slow or unexplained, and the job is timeline capture, per-kernel metrics, roofline reading, occupancy analysis, or NVTX annotation. |
| Authority | Reversible local. Writes are limited to profiler reports and logs under the project tree; rollback is deleting those files. No remote mutation. |
| Side effect | Report files (.nsys-rep, .ncu-rep, CSV), possibly elevated-privilege profiling runs, and a diagnosis. |
| Done | The bottleneck is named with a measured metric, and the recommended fix targets that measurement. |
Inputs
- Workload (required): the application or kernel to profile and one representative input.
- Question (required): timeline overlap, per-kernel cost, occupancy, or memory versus compute balance.
- GPU and toolkit (required):
nvidia-smiandnvcc --version. Grounded current stable: CUDA Toolkit 13.3 Update 1. - Profiling permissions (required on locked-down hosts): membership checks in step 7.
Procedure
- Pick the tool by question. System-wide timeline across CPU, GPU, and CUDA API calls goes to Nsight Systems (
nsys). Per-kernel counters and stall analysis goes to Nsight Compute (ncu). A single metric in CI goes toncu --metrics. Done when: the tool matches the question. - Build with line info and no
-G:
nvcc -lineinfo -O3 -arch=sm_80 -o app main.cu
Done when: the profiled binary carries line info for source correlation. 3. Capture the timeline first:
nsys profile --trace=cuda,nvtx,osrt --output=timeline ./app
nsys stats timeline.nsys-rep # CLI summary
nsys-ui timeline.nsys-rep # GUI
Read the timeline for gaps between launches (CPU bottleneck or sync stalls), cudaDeviceSynchronize waits, overlap between copies and kernels across streams, and CUDA API overhead. --capture-range=cudaProfilerApi limits capture to a marked region. Done when: each observed gap is attributed to a named cause.
4. Annotate phases with NVTX so timeline bands match application stages:
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 120 lines · 42 tokens per session scan A 0535cd60e9ff
cuda-profiling is a skill published in the GitHub repository OutlineDriven/outline-driven-development (52 stars, last pushed 3d ago), licensed Apache-2.0. It adds 42 tokens to every session and 1,611 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-06.
Other skills, from other repositories
audit-project
Run an iterative multi-agent code audit until critical and high findings are resolved. Use when the user says "audit my code", "find all the bugs", "deep code audit", "iterative review", or "review until clean".
debug
Hypothesis-driven debugging. Use when a test fails, a crash or exception occurs, output is wrong, or an intermittent flake has no obvious cause.
optimize
Use when asked to optimize code, speed up a path, reduce allocations, repair a regression, or profile a target. Not for remote, credential, publish, deploy, or irreversible changes.
ios-build-fix
Use when asked to run /ios-build-fix to fix a failing iOS build, regenerate an Xcode project from project.yml, or correct UI behavior. Not for a clean rebuild: use ios-build-cleanup.
strike-the-root
Use when a bug, failure, flake, regression, review finding, or ticket needs the core fixed so it cannot recur. Not for greenfield features: use tdd. Not for style-only review or typo-class one-liners.
extremely-optimize
Use when asked to run a performance campaign against a measured floor. Not for hypothesis-only analysis without mutation: use fastopt.