Borrowing it
Nothing to install: this file belongs to maoxx241/vllm-ascend-workspace. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/maoxx241/vllm-ascend-workspace/main/.agents/skills/ascend-triton-workflow/SKILL.mdgit clone --depth 1 https://github.com/maoxx241/vllm-ascend-workspaceWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/maoxx241/vllm-ascend-workspace/ascend-triton-workflow)<a href="https://agentmods.dev/skills/maoxx241/vllm-ascend-workspace/ascend-triton-workflow"><img src="https://agentmods.dev/badge/skills/maoxx241/vllm-ascend-workspace/ascend-triton-workflow/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/maoxx241/vllm-ascend-workspace/ascend-triton-workflow"><img src="https://agentmods.dev/badge/skills/maoxx241/vllm-ascend-workspace/ascend-triton-workflow.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00093 | $0.00569 |
| Opus 5 | $0.00046 | $0.00284 |
| Sonnet 5 | $0.00019 | $0.00114 |
| Haiku 4.5 | $0.00009 | $0.00057 |
Grade A, and why
ascend-triton-workflow scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 45 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Ascend Triton Workflow
Coordinate the lifecycle without duplicating the implementation owned by the stage Skills.
Workflow
- Freeze the source, operator contract, target SoC, software versions, case set, and requested performance objective.
- Query
.agents/knowledge/for relevant capability, validation, and failure-signature facts. Treat missing facts as unknown. - Run
scripts/triton_workflow.py planto create the stage plan and parent Run Manifest. - Execute required stages with their owners:
- first correct kernel or GPU migration:
ascend-triton-operator-development; - compile and full-case correctness gate:
ascend-triton-kernel-validation; - single-kernel profiling and performance iteration:
ascend-triton-kernel-optimization.
- first correct kernel or GPU migration:
- Before remote execution, establish
remote-code-parity; runtorch_npuand Triton only in a managed Ascend environment. - Link every child Run Manifest to its planned stage with
link. - Run
finalizeand deliver the workflow report with missing, failed, and untested coverage explicit.
Entry point
scripts/triton_workflow.py provides:
plan: validate the workflow config, create ordered stage items, and initialize a parent Run Manifest;link: bind one terminal child Run Manifest to one stage without overwriting prior evidence;finalize: aggregate required-stage evidence and produceworkflow-report.md.
Read:
- Behavior contract for config, stage, link, and status semantics.
- Command recipes for the complete lifecycle.
- Acceptance before claiming an operator workflow complete.
Rules
- Keep stage ownership strict; do not implement migration, validation, or optimization inside this Skill.
- Never let performance evidence substitute for correctness evidence.
- A passed workflow requires every required stage to have a passed terminal child manifest.
- A failed child makes the workflow failed; missing, unsupported, cancelled, or inconclusive required evidence makes it inconclusive.
- Record GPU timing only as cross-platform context; use a comparable NPU baseline for NPU optimization acceptance.
- Keep orchestration state under
.vaws-local/ascend-triton/workflows/.
What ships with it
6 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 45 lines · 93 tokens per session scan A 0e3d06b86eb3
ascend-triton-workflow is a skill published in the GitHub repository maoxx241/vllm-ascend-workspace (36 stars, last pushed 7d ago), licensed MIT. It adds 93 tokens to every session and 569 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
jetson-video-pipeline
Use when executing and verifying Jetson Video Codec SDK or PyNvVideoCodec encode/decode, transcode, segmentation, container decode, AV1, or acceptance workflows with exact artifact handoffs.
jetson-validate-image
Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.
holohub-app-lifecycle
Use for non-failing HoloHub app work with ./holohub: scaffold, build, run, test, visual evidence, lint, and flow benchmarking.
code-plan
Turn a task description and repository into a structured implementation plan (files to create, files to modify, tests to add, risks).
ros2-robotics
Best practices for ROS 2 robotics development, covering package structure, nodes, topics/services/actions, launch files, QoS, tf2 transforms, and testing. Use when creating ROS 2 packages, writing nodes in rclpy or rclcpp, defining custom messages/services/actions, writing launch files, configuring QoS profiles…
SmartHome Video Anomaly Benchmark
VLM evaluation suite for video anomaly detection in smart home camera footage.