Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add nebius/nebius-physical-ai --skill gpu-cluster-provisioninggit clone --depth 1 https://github.com/nebius/nebius-physical-aiWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/nebius/nebius-physical-ai/gpu-cluster-provisioning)<a href="https://agentmods.dev/skills/nebius/nebius-physical-ai/gpu-cluster-provisioning"><img src="https://agentmods.dev/badge/skills/nebius/nebius-physical-ai/gpu-cluster-provisioning/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/nebius/nebius-physical-ai/gpu-cluster-provisioning"><img src="https://agentmods.dev/badge/skills/nebius/nebius-physical-ai/gpu-cluster-provisioning.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 1 finding, up to high
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- high Tool Misuse · line 47 Tool parameters are crafted to achieve unintended or unsafe behavior. Parameter abuse can bypass intended safety checks (e.g. shell=True, --force, dangerous glob patterns).Fix: Validate all tool parameters against an allowlist. Reject dangerous parameter values (shell=True, --force, -rf /) and use safe defaults.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00067 | $0.02343 |
| Opus 5 | $0.00034 | $0.01171 |
| Sonnet 5 | $0.00013 | $0.00469 |
| Haiku 4.5 | $0.00007 | $0.00234 |
Grade A, and why
gpu-cluster-provisioning scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 197 lines — stays where its author put it; the contents beside it link to each section on GitHub.
GPU cluster provisioning and driver strategy
Getting GPU nodes is the easy part. Getting nodes whose GPUs actually work — and
knowing before you submit a job that they do — is what this skill covers. It
complements skills/atomic/gpu-selection/SKILL.md (which GPU to ask for) and
skills/tools/nebius-infra/SKILL.md (config, storage, teardown) with the driver
and readiness decisions made at provisioning time.
Source of record: docs/workbench/mk8s-gpu-driver-strategy.md.
The driver decision: default to managed, do not reach for operator
One policy covers both direct npa cluster provisioning and npa fleet: GPU
node groups use a Nebius managed-driver image by default, and CPU-only node
groups get no GPU driver settings at all.
--gpu-driver-mode |
Behavior | Use when |
|---|---|---|
auto (default) |
Managed-driver image whenever the recipe/provider supports it | Almost always |
managed-image |
Explicitly require the managed-driver image | Pinning the operational contract |
operator |
In-cluster NVIDIA GPU Operator driver path | Supported diagnostics or a recipe that requires it |
npa cluster up --gpu-driver-mode auto --managed-driver-preset cuda13.0
The default managed preset is cuda13.0. The same values exist on
npa provision-if-absent, and in a fleet spec under defaults or a single
cluster (gpu_driver_mode, managed_driver_preset).
Operator mode is unsafe on NVSwitch topologies. Fabric Manager can start before Network Operator/MOFED exposes the host InfiniBand management devices, and the result is nodes that exist but cannot run CUDA. It presents as:
NebiusGPUError=True- Fabric status
In ProgressorN/A - CUDA
system not yet initialized - Fabric Manager / NVLSM
umad_open_port()orIB_ERRORfailures
NPA identifies an NVSwitch-risk topology from an explicit GPU cluster or a
multi-GPU SXM/NVL preset and rejects operator mode unless you pass
--allow-unsafe-nvswitch-operator. That flag is a diagnostics acknowledgement,
not a workaround: if a deploy fails and the suggested fix is that flag, the fix
is almost always auto instead.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago Changed · +4 lines edf2d42b6c58
- 6d ago Changed · +19 lines 3efa36208d19
- 10d ago First seen · 174 lines · 67 tokens per session scan A 3d6509054796
gpu-cluster-provisioning is a skill published in the GitHub repository nebius/nebius-physical-ai (27 stars, last pushed yesterday), licensed Apache-2.0. It adds 67 tokens to every session and 2,343 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
azd-deployment
Deploy containerized frontend + backend applications to Azure Container Apps with remote builds, managed identity, and idempotent infrastructure.
openshell-cli
Guide agents through using the OpenShell CLI (openshell) for sandbox management, gateway registration, provider configuration and refresh, policy iteration, settings, service exposure, BYOC workflows, and inference routing. Covers basic through advanced multi-step workflows. Trigger keywords - openshell, sandbox…
langbot-deploy
Deploy and configure a LangBot instance — Docker / Docker Compose, Kubernetes, the config.yaml model, the Box sandbox runtime, the plugin runtime, and the global API key. Use when installing, deploying, upgrading, or configuring LangBot in production or self-hosted environments. Triggers on "deploy langbot", "langbot…
compute-env-setup
Set up a compute environment on a remote provider so Claude Science jobs can run there. Covers direct SSH/conda hosts, Slurm clusters, container-via-bridge runners, and managed-API providers (Modal, GCP, RunPod). Use when standing up a new provider, porting an env to a different backend, adding a tool that needs its…
azure-cloud-migrate
Assess and migrate cross-cloud workloads to Azure with reports and code conversion. Supports Lambda→Functions, Beanstalk/Heroku/App Engine→App Service, Fargate/Kubernetes/Cloud Run/Spring Boot→Container Apps. WHEN: migrate Lambda to Functions, AWS to Azure, migrate Beanstalk, migrate Heroku, migrate App Engine, Cloud…
atmos-helmfile
Helmfile orchestration: sync/apply/destroy/diff, Kubernetes deployments, varfile generation, EKS integration, source management.