Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add nebius/nebius-physical-ai --skill golden-evalgit clone --depth 1 https://github.com/nebius/nebius-physical-aiWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/nebius/nebius-physical-ai/golden-eval)<a href="https://agentmods.dev/skills/nebius/nebius-physical-ai/golden-eval"><img src="https://agentmods.dev/badge/skills/nebius/nebius-physical-ai/golden-eval/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/nebius/nebius-physical-ai/golden-eval"><img src="https://agentmods.dev/badge/skills/nebius/nebius-physical-ai/golden-eval.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00046 | $0.01363 |
| Opus 5 | $0.00023 | $0.00681 |
| Sonnet 5 | $0.00009 | $0.00273 |
| Haiku 4.5 | $0.00005 | $0.00136 |
Grade A, and why
golden-eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 122 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Golden eval (does this container actually work?)
A golden eval is the minimal tested rerun that proves one container image is functional. It answers a narrow, valuable question — does this image run its own smoke on a real GPU? — without standing up a cluster or a workflow. Reach for it after building or republishing an image, and when triaging whether a failure is the image or the pipeline around it.
The manifest is npa/src/npa/smoke/golden_evals.yaml
(format npa_golden_evals_v1); every container in CONTAINER_IMAGE_NAMES must
have an entry, enforced by npa/tests/smoke/test_golden_eval_manifest.py.
Inspect before running
npa workbench golden-eval list # every container, kind, gpu, status
npa workbench golden-eval list --output json
npa workbench golden-eval show <container> # full safety + Physical AI record
Read the status column before spending anything:
ready— runnable now.gpu-gated— needs a real GPU; a local run without one proves nothing.blocked-on-upstream— excluded from batch runs unless you pass--include-blocked. A failure here is expected and is not your regression.needs-image-update— the manifest and the published image disagree; rebuild before drawing conclusions.
kind tells you what is actually exercised: container-smoke, server-smoke,
entrypoint-smoke, workflow-smoke, or build-import. A build-import passing
means the package imports, not that the tool works. gpu is required,
optional, or none.
Three execution tiers, cheapest first
npa workbench golden-eval run <container> # dry run (default)
npa workbench golden-eval run <container> --execute # local runtime
npa workbench golden-eval run <container> --serverless # one GPU, real image
npa workbench golden-eval run <container> --serverless --gpu h200 --timeout 40m
Dry run is the default and prints the command. Use it to see exactly what would execute — often enough to answer a question without running anything.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 122 lines · 46 tokens per session scan A d9c3e28cc82a
golden-eval is a skill published in the GitHub repository nebius/nebius-physical-ai (27 stars, last pushed today), licensed Apache-2.0. It adds 46 tokens to every session and 1,363 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
verify
Build the computer-use-demo image and drive the Streamlit UI headlessly to verify changes end-to-end.
test-locally
Build and deploy a full local OpenMetadata stack with Docker to test your connector in the UI. Handles code generation, build optimization, health checks, and guided testing.
redamon-testing
How RedAmon tests actually run and how to author them: the per-file Docker gate, the unit/integration/live tiers, and the failure modes that make a green run a lie. Trigger: editing any test.py, .test.ts(x) or tests/.sh; a test that is red, skipped or xfailed; a request to "run the tests", "make it green" or check…
zero-script-qa
Zero Script QA — test without scripts using structured JSON logging and Docker monitoring. Triggers: zero-script-qa, log testing, docker logs, QA.
local-test
Build, run, and test IronClaw locally using Docker containers and Chrome MCP browser automation.
testcontainers-integration-tests
Write integration tests using TestContainers for .NET with xUnit. Covers infrastructure testing with real databases, message queues, and caches in Docker containers instead of mocks.