Borrowing it
Nothing to install: this file belongs to maoxx241/vllm-ascend-workspace. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/maoxx241/vllm-ascend-workspace/main/.agents/skills/ascend-triton-kernel-validation/SKILL.mdgit clone --depth 1 https://github.com/maoxx241/vllm-ascend-workspaceWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/maoxx241/vllm-ascend-workspace/ascend-triton-kernel-validation)<a href="https://agentmods.dev/skills/maoxx241/vllm-ascend-workspace/ascend-triton-kernel-validation"><img src="https://agentmods.dev/badge/skills/maoxx241/vllm-ascend-workspace/ascend-triton-kernel-validation/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/maoxx241/vllm-ascend-workspace/ascend-triton-kernel-validation"><img src="https://agentmods.dev/badge/skills/maoxx241/vllm-ascend-workspace/ascend-triton-kernel-validation.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00122 | $0.00623 |
| Opus 5 | $0.00061 | $0.00311 |
| Sonnet 5 | $0.00024 | $0.00125 |
| Haiku 4.5 | $0.00012 | $0.00062 |
Grade A, and why
ascend-triton-kernel-validation scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 43 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Ascend Triton Kernel Validation
Turn a candidate into explicit compile and correctness evidence. A successful process exit or benchmark run is not correctness proof.
Workflow
- Freeze the candidate hash, trusted reference, predeclared tolerances, target environment, and explicit case matrix.
- Query
.agents/knowledge/with any observed compile, runtime, or numerical signature before repeating diagnosis. - Run
scripts/triton_validation.py plan. It invokesvalidate_triton_impl.pyand rejects missing Triton kernels, aModelNew.forwardpath that does not launch them, or reachable PyTorch tensor computation fallback. - Before remote execution, establish
remote-code-parity. Run every planned case on a managed Ascend NPU; do not runtorch_npulocally. - Compare shape and dtype first, then NaN/Inf behavior and numeric values. Preserve raw stdout, stderr, stack, and comparison artifacts.
- Normalize one result per case and run
record. Do not overwrite evidence. - Run
analyze. Onlypassed_cases == total_cases > 0produces a passed manifest. - Hand the passed manifest to development or optimization. Do not benchmark a failed or inconclusive candidate.
Entry points
scripts/validate_triton_impl.py: AST-only fallback and launch gate.scripts/triton_validation.py:plan: validate config and candidate, create case matrix and Run Manifest v1;record: accept one normalized case result;analyze: classify the full matrix and generate the report.
Read:
- Behavior contract for config, result, and status schemas.
- Case design before choosing the matrix and tolerances.
- Command recipes for the lifecycle.
- Acceptance before declaring correctness.
Rules
- Keep compile error, runtime error, numerical mismatch, unsupported, and missing evidence distinct.
- Never silently cast, make inputs contiguous, remove difficult shapes, or relax tolerance to pass.
- Load masks protect readable input addresses; store masks protect writable output positions.
- Treat tail blocks, fully masked blocks, non-power-of-two shapes, and dynamic specialization boundaries as first-class cases.
- Keep run state under
.vaws-local/ascend-triton/validation/.
What ships with it
8 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 43 lines · 122 tokens per session scan A 185ce383a3d1
ascend-triton-kernel-validation is a skill published in the GitHub repository maoxx241/vllm-ascend-workspace (36 stars, last pushed 7d ago), licensed MIT. It adds 122 tokens to every session and 623 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
jetson-video-pipeline
Use when executing and verifying Jetson Video Codec SDK or PyNvVideoCodec encode/decode, transcode, segmentation, container decode, AV1, or acceptance workflows with exact artifact handoffs.
jetson-validate-image
Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.
holohub-app-lifecycle
Use for non-failing HoloHub app work with ./holohub: scaffold, build, run, test, visual evidence, lint, and flow benchmarking.
code-plan
Turn a task description and repository into a structured implementation plan (files to create, files to modify, tests to add, risks).
ros2-robotics
Best practices for ROS 2 robotics development, covering package structure, nodes, topics/services/actions, launch files, QoS, tf2 transforms, and testing. Use when creating ROS 2 packages, writing nodes in rclpy or rclcpp, defining custom messages/services/actions, writing launch files, configuring QoS profiles…
SmartHome Video Anomaly Benchmark
VLM evaluation suite for video anomaly detection in smart home camera footage.