vllm-ascend-workspace: Skill for Codex

.agents/skills/npu-fleet-monitor/SKILL.md

npu-fleet-monitor is a skill for Codex from maoxx241/vllm-ascend-workspace. It costs 71 tokens per session (639 once invoked), scanned A, original, MIT.

A bootstrap and command-line entry point for vaws-top, a tool that discovers and inspects a fleet of NPU servers. NPUs are processors commonly used for machine-learning workloads.

In plain words
What is it for?
Use it to deploy or manage the local monitor, discover available servers, check capacity, inspect a host, view mounted storage, or see NPU process details.
Why use it?
It helps locate the fleet-monitoring tool and provides basic ways to see servers, capacity, status, mounts, and running processes.

Skill for Codex

Written for Codex: agents/openai.yaml present. Also seen: installed under .agents/ (shared by several agents).

This is maoxx241/vllm-ascend-workspace's own configuration. It tells Codex how to work on vllm-ascend-workspace itself, so it is not a mod to install elsewhere. Copy it as a starting point and replace the rules that are about this project. Everything vllm-ascend-workspace configures →

Reuse

Borrowing it

Nothing to install: this file belongs to maoxx241/vllm-ascend-workspace. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.

Copy the file
curl -O https://raw.githubusercontent.com/maoxx241/vllm-ascend-workspace/main/.agents/skills/npu-fleet-monitor/SKILL.md
Clone the repo
git clone --depth 1 https://github.com/maoxx241/vllm-ascend-workspace

Made for: Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for npu-fleet-monitor

README.md
[![agentmods](https://agentmods.dev/badge/skills/maoxx241/vllm-ascend-workspace/npu-fleet-monitor/github.svg)](https://agentmods.dev/skills/maoxx241/vllm-ascend-workspace/npu-fleet-monitor)
Your own site
<a href="https://agentmods.dev/skills/maoxx241/vllm-ascend-workspace/npu-fleet-monitor"><img src="https://agentmods.dev/badge/skills/maoxx241/vllm-ascend-workspace/npu-fleet-monitor/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for npu-fleet-monitor

Your own site · 80×15
<a href="https://agentmods.dev/skills/maoxx241/vllm-ascend-workspace/npu-fleet-monitor"><img src="https://agentmods.dev/badge/skills/maoxx241/vllm-ascend-workspace/npu-fleet-monitor.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 71 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 639 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • NVIDIA SkillSpector pass 7 Sept 2026
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00071 $0.00639
Opus 5 $0.00036 $0.00319
Sonnet 5 $0.00014 $0.00128
Haiku 4.5 $0.00007 $0.00064

Measured 12d ago against content hash b3abb444b248, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-12, from the pricing page.

Security

Grade A, and why

npu-fleet-monitor scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.

The scan reads SKILL.md. This mod also ships 2 executable files (scripts/manage_monitor.py, tests/test_manage_monitor.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agents/skills/npu-fleet-monitor/SKILL.md · 55 lines

How it starts

The opening of the file, as written. The whole thing — 55 lines — stays where its author put it; the contents beside it link to each section on GitHub.

vaws-top entry

Keep the application and its complete Agent instructions on the standalone vaws-top branch. This main-branch Skill is only the bootstrap entry.

Run the helper on the host execution plane. Deploy or reconcile:

python3 .agents/skills/npu-fleet-monitor/scripts/manage_monitor.py ensure

Locate, inspect, restart, or stop:

python3 .agents/skills/npu-fleet-monitor/scripts/manage_monitor.py status
python3 .agents/skills/npu-fleet-monitor/scripts/manage_monitor.py restart
python3 .agents/skills/npu-fleet-monitor/scripts/manage_monitor.py stop

The final JSON includes worktree and agent_skill. Use the returned worktree as <vaws-top> below.

Basic CLI

python3 <vaws-top>/scripts/vaws-top.py servers
python3 <vaws-top>/scripts/vaws-top.py capacity --min-idle 4 --max-age 180
python3 <vaws-top>/scripts/vaws-top.py status HOST
python3 <vaws-top>/scripts/vaws-top.py status HOST --cache
python3 <vaws-top>/scripts/vaws-top.py mounts HOST
python3 <vaws-top>/scripts/vaws-top.py --json npu HOST --process-details

status HOST is live by default; add --cache when stored data is sufficient. servers, capacity, mounts, and npu use cached observations by default; commands that support it accept --live. Add --json for structured output. A live query asks the centralized service to probe once; do not follow a successful result with ad hoc SSH. Capacity is observed availability, not a reservation.

Basic MCP

Run the stdio server directly or register it in the Agent's MCP configuration:

[mcp_servers.vaws_top]
command = "python3"
args = ["<vaws-top>/scripts/vaws-top-mcp.py"]
env = { VAWS_TOP_URL = "http://127.0.0.1:8789" }

The basic tools are list_npu_servers, find_npu_capacity, server_status, npu_status, and list_mounts. Host-query tools default to cached data in MCP; pass mode="live" for a fresh centralized probe.

Before advanced fleet selection, process attribution, mount discovery, or operational changes, read the returned agent_skill completely and follow it.

Read the full file on GitHub · 55 lines

Files

What ships with it

3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 12d ago First seen · 55 lines · 71 tokens per session scan A b3abb444b248

Subscribe to this mod's changes

npu-fleet-monitor is a skill published in the GitHub repository maoxx241/vllm-ascend-workspace (36 stars, last pushed 8d ago), licensed MIT. It adds 71 tokens to every session and 639 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

gke-compute-classes

Configures, optimizes, and troubleshoots GKE ComputeClasses. Use when configuring Spot VMs with on-demand fallback, targeting specific accelerators (GPUs/TPUs) or machine families, restricting ComputeClass access, or debugging pending pods related to node pool auto-creation. Do not use for cluster-level Node Auto…

google/skills · 83 tokens

jetson-diagnostic

Read-only Jetson health snapshot for identity, memory, GPU, thermal, power, storage, services, and top processes.

NVIDIA/skills · 30 tokens

doca-socket-relay

Use this skill when the operator is driving the DOCA Socket Relay to bridge a socket-oriented host application onto a BlueField DPU peer without rewriting it — picking the deployment shape (in-process, sidecar, or BlueField service container), configuring the host-side socket and the DPU-side forwarding endpoint…

NVIDIA/skills · 236 tokens

offensive-z-wave

Z-Wave attack methodology — sniffing with Z-Force / EZ-Wave / RTL-SDR + ZniffMobile, S0 (legacy) network-key derivation flaw and key reuse, S2 (modern) ECDH commissioning analysis, replay/injection on unauthenticated nodes, default-key brute-force on test deployments, and home-automation hub pivots. Use when targeting…

SnailSploit/Claude-Red · 113 tokens

hsb-flash

Flash the FPGA on an HSB board connected to an NVIDIA devkit. Supports HSB Lattice boards (FPGA versions 2407, 2412, 2507, 2510) and Leopard Imaging VB1940 "all-in-one" cameras (FPGA versions 2507, 2510). Uses release-specific YAML manifests and board-type-specific program commands. Lattice and VB1940 commands must…

NVIDIA/skills · 94 tokens

jetson-validate-image

Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.

NVIDIA/skills · 50 tokens