Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add aws-neuron/neuron-agentic-development --skill neuron-nki-profile-queryinggit clone --depth 1 https://github.com/aws-neuron/neuron-agentic-developmentWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/aws-neuron/neuron-agentic-development/neuron-nki-profile-querying)<a href="https://agentmods.dev/skills/aws-neuron/neuron-agentic-development/neuron-nki-profile-querying"><img src="https://agentmods.dev/badge/skills/aws-neuron/neuron-agentic-development/neuron-nki-profile-querying.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00207 | $0.04370 |
| Opus 5 | $0.00103 | $0.02185 |
| Sonnet 5 | $0.00041 | $0.00874 |
| Haiku 4.5 | $0.00021 | $0.00437 |
Grade A, and why
neuron-nki-profile-querying scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
on localhost. No deployment, no remote service — just the CLI and curl. How it starts
The opening of the file, as written. The whole thing — 438 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Profile Querying
Run SQL queries against NKI kernel profile data using neuron-explorer view.
This ingests NEFF+NTFF into parquet and exposes a DuckDB-backed API server
on localhost. No deployment, no remote service — just the CLI and curl.
For more advanced analysis, use python on parquet to compute performance bounds and investigate precise inefficiencies within arbitrary execution intervals.
What you need: A compiled NEFF file and a captured NTFF trace file.
These come from /neuron-nki-profiling or from running a kernel with the right
env vars and neuron-explorer capture.
Quick Start
# Ingest and start API server (no web UI)
neuron-explorer view \
-n ./kernel.neff \
-s ./profile.ntff \
--data-path ~/.local/share/neuron-profile \
--display-name my-kernel \
--disable-ui &
# Wait for server
sleep 10
# Query
curl -s -X POST http://localhost:3002/api/v1/db/my-kernel/_search \
-H 'Content-Type: application/json' \
-d '{"type":"databaseExplorerQuery","tableName":"Summary","query":"SELECT total_time, mfu_estimated_percent, tensor_engine_active_time_percent, dma_active_time_percent FROM Summary"}'
That's it. Ingest, serve, query.
Prerequisites
neuron-explorerinstalled (comes with AL2023 DLAMI oraws-neuronx-tools)- NEFF file (compiled kernel binary) + NTFF file (execution trace)
Check availability:
which neuron-explorer && neuron-explorer --version
If not found, check /opt/aws/neuron/bin/neuron-explorer.
Step-by-Step Workflow
Step 0: Check Profile Quality (Re-profile if Needed)
Note: This step is specific to NKI kernel development. If you are querying a profile that was generated outside of an NKI workflow, skip to Step 1.
Disclaimer: Query results are only as good as the profile. If the NEFF/NTFF were captured without the right env vars, key tables (DmaPacket, DmaPacketAggregated) may be empty and source-level attribution will be missing.
Check whether the profile has the data you need:
# After ingesting (Step 2), check for DMA packet data
curl -s -X POST http://localhost:3002/api/v1/db/${PROFILE_NAME}/_search \
-H 'Content-Type: application/json' \
-d '{"type":"databaseExplorerQuery","tableName":"DmaPacket","query":"SELECT COUNT(*) as cnt FROM DmaPacket"}'
What ships with it
34 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- references/example-bounds-analysis.md 13 KB
- references/getting-started.md 3.1 KB
- references/investigations/dma_efficiency.md 11 KB
- references/investigations/redundant_dma_transfers.md 10 KB
- references/investigations/redundant_te_transposes.md 7.1 KB
- references/investigations/te_inefficiency.md 7.6 KB
- references/performance-bounds.md 17 KB
- references/schema/ActiveTime.yaml 1.4 KB
- references/schema/DmaPacket.yaml 4.9 KB
- references/schema/DmaPacketAggregated.yaml 8.2 KB
- references/schema/DmaQueuesInfo.yaml 1.4 KB
- references/schema/DmaUsage.yaml 1.1 KB
- references/schema/Error.yaml 1.4 KB
- references/schema/Flow.yaml 1.0 KB
- references/schema/HbmUsage.yaml 1.1 KB
- references/schema/HbmUsageSummaryByType.yaml 817 B
- references/schema/HostMemUsage.yaml 1.3 KB
- references/schema/Instruction.yaml 13 KB
- references/schema/KernelInstructions.yaml 2.4 KB
- references/schema/KernelIterationVariables.yaml 1.7 KB
- references/schema/KernelStackFrames.yaml 1.5 KB
- references/schema/Metadata.yaml 8.3 KB
- references/schema/NeffHeader.yaml 2.0 KB
- references/schema/PendingDma.yaml 706 B
- references/schema/PsumUsage.yaml 2.2 KB
- references/schema/SbufAllocation.yaml 3.2 KB
- references/schema/SbufUsage.yaml 1.5 KB
- references/schema/SchemaFields.yaml 2.3 KB
- references/schema/SemaphoreUpdate.yaml 871 B
- references/schema/Summary.yaml 30 KB
- references/schema/TensorInfo.yaml 2.1 KB
- references/schema/Throttle.yaml 1.3 KB
- references/schema/ThrottleSummary.yaml 1.5 KB
- references/schema/Warning.yaml 595 B
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 7d ago First seen · 438 lines · 207 tokens per session scan A 46eeb97faff6
neuron-nki-profile-querying is a skill published in the GitHub repository aws-neuron/neuron-agentic-development (57 stars, last pushed 18d ago), licensed Apache-2.0. It adds 207 tokens to every session and 4,370 once invoked, about $0.0010 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
gke-compute-classes
Configures, optimizes, and troubleshoots GKE ComputeClasses. Use when configuring Spot VMs with on-demand fallback, targeting specific accelerators (GPUs/TPUs) or machine families, restricting ComputeClass access, or debugging pending pods related to node pool auto-creation. Do not use for cluster-level Node Auto…
jetson-diagnostic
Read-only Jetson health snapshot for identity, memory, GPU, thermal, power, storage, services, and top processes.
doca-socket-relay
Use this skill when the operator is driving the DOCA Socket Relay to bridge a socket-oriented host application onto a BlueField DPU peer without rewriting it — picking the deployment shape (in-process, sidecar, or BlueField service container), configuring the host-side socket and the DPU-side forwarding endpoint…
offensive-z-wave
Z-Wave attack methodology — sniffing with Z-Force / EZ-Wave / RTL-SDR + ZniffMobile, S0 (legacy) network-key derivation flaw and key reuse, S2 (modern) ECDH commissioning analysis, replay/injection on unauthenticated nodes, default-key brute-force on test deployments, and home-automation hub pivots. Use when targeting…
hsb-flash
Flash the FPGA on an HSB board connected to an NVIDIA devkit. Supports HSB Lattice boards (FPGA versions 2407, 2412, 2507, 2510) and Leopard Imaging VB1940 "all-in-one" cameras (FPGA versions 2507, 2510). Uses release-specific YAML manifests and board-type-specific program commands. Lattice and VB1940 commands must…
jetson-validate-image
Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.