Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/intelligent-internet/zenith/benchmark-validatornpx skills add Intelligent-Internet/zenith --skill benchmark-validatorgit clone --depth 1 https://github.com/Intelligent-Internet/zenithWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00044 | $0.01454 |
| Opus 5 | $0.00022 | $0.00727 |
| Sonnet 5 | $0.00009 | $0.00291 |
| Haiku 4.5 | $0.00004 | $0.00145 |
Grade A, and why
benchmark-validator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 153 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Benchmark Validator
Use this skill for a validation assignment targeting exactly one benchmark-related assertion.
For optimization EXP-* targets, you are not selecting the winner or optimizing
further. You independently decide whether the experiment produced a trustworthy
outcome under its contract.
For optimization EXP-* targets, act as an adversarial tester for the selected
candidate. The contract is the minimum bar, not the whole test plan. Before
promotion, define and run a compact but comprehensive validation plan that
attacks the candidate's likely correctness and performance failure modes.
For engineering VAL-* or legacy engineering targets, do not use optimization
outcome semantics. Pass only when the assigned benchmark, performance,
correctness, and guardrail behavior required by the assignment and contract is
proven with fresh evidence.
Inputs
Read:
- Validation assignment.
- The single assigned benchmark-related contract.
AGENTS.md.- Experiment ledger path cited by the assignment or contract.
- Candidate artifact/ref/patch/checksum cited by the assignment, contract, or ledger.
- Measurement protocol, correctness/guardrail commands, protected files, and source/baseline refs cited by the contract.
- For optimization targets, also inspect visible runtime resources read-only when useful: changed files, benchmark scripts, fixtures, public outputs, baseline/reference artifacts, verifier-adjacent code, logs, and generated artifacts. Do not mutate these resources, copy protected artifacts into a submission, or make submitted code depend on evaluator-only locations.
If the assignment targets more than one benchmark-related assertion, fail the assignment as too broad and request attention.
Procedure
- Identify artifacts
- Parent/baseline ref.
- Candidate ref or patch artifact.
- Benchmark command, run count, aggregation, variance policy.
- Correctness and guardrail commands.
- Protected benchmark/scoring/data/verifier files.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 153 lines · 44 tokens per session scan A bf618457802d
benchmark-validator is a skill published in the GitHub repository Intelligent-Internet/zenith (283 stars, last pushed 26d ago), licensed Apache-2.0. It adds 44 tokens to every session and 1,454 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
agent-harness-fault-injection
Use when an agent workflow needs deterministic recovery evidence for sandbox, MCP/tool, worker, checkpoint, memory, or orchestration failures.
terraform-skill
Use when working with Terraform or OpenTofu - creating modules, writing tests (native test framework, Terratest), setting up CI/CD pipelines, reviewing configurations, choosing between testing approaches, debugging state issues, implementing security scanning (trivy, checkov), or making infrastructure-as-code…
xlsx
当电子表格文件是主要输入或输出时使用此技能。这意味着用户想要:打开、读取、编辑或修复现有的 .xlsx、.xlsm、.csv 或 .tsv 文件(例如添加列、计算公式、格式化、制图、清理混乱数据);从头创建新的电子表格或从其他数据源创建;或在表格文件格式之间进行转换。当用户通过名称或路径引用电子表格文件时特别触发——即使是随意提及(如"我下载目录里的 xlsx")——并且想对其进行操作或从中生成内容。也适用于将混乱的表格数据文件(格式错误的行、错位的表头、垃圾数据)清理或重构为规范的电子表格。交付物必须是电子表格文件。当主要交付物是 Word 文档、HTML 报告、独立 Python 脚本、数据库管道或 Google Sheets…
terraform-cli-setup
Terraform CLI 安装与初始化技能。当用户本地未安装 Terraform 时自动完成安装,确保 terraform 命令可用并能执行 init/validate。不负责 Provider 凭证配置,凭证在实际使用时由 terraform-skill 引导。.
pr-integration-test
Design, implement, and validate Intelligent Terminal integration tests for a target pull request or regression. Use when asked to add PR integration tests, convert a bug fix into E2E coverage, prove existing behavior still works, map tests to the release checklist, or verify E2E reports mark checklist cases complete.
oma-scholar
Scholarly research companion using Knows sidecar spec (.knows.yaml). Generates, validates, reviews, queries, and compares structured research-paper sidecars, and fetches them from knows.academy. Use for academic literature search, survey synthesis, paper authoring assistance, and peer review with token-efficient…