benchmark-validator

A procedure for checking whether a benchmark result can be trusted. A benchmark is a measured test of performance or correctness, and the procedure checks the assigned claim against fresh evidence and guardrails.

In plain words
What is it for?
Use it to validate optimization experiments, performance claims, correctness checks, and safety conditions for engineering benchmark targets.
Why use it?
It helps catch misleading, incomplete, or incorrect results before an implementation or experiment is accepted.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/intelligent-internet/zenith/benchmark-validator
Any agent
npx skills add Intelligent-Internet/zenith --skill benchmark-validator
Clone the repo
git clone --depth 1 https://github.com/Intelligent-Internet/zenith

Made for: Claude Code, Codex.

Per session 44 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,454 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00044 $0.01454
Opus 5 $0.00022 $0.00727
Sonnet 5 $0.00009 $0.00291
Haiku 4.5 $0.00004 $0.00145

Measured 3d ago against content hash bf618457802d, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

benchmark-validator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

zenith/src/zenith_harness/bundled/skills/benchmark-validator/SKILL.md · 153 lines

How it starts

The opening of the file, as written. The whole thing — 153 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Benchmark Validator

Use this skill for a validation assignment targeting exactly one benchmark-related assertion.

For optimization EXP-* targets, you are not selecting the winner or optimizing further. You independently decide whether the experiment produced a trustworthy outcome under its contract.

For optimization EXP-* targets, act as an adversarial tester for the selected candidate. The contract is the minimum bar, not the whole test plan. Before promotion, define and run a compact but comprehensive validation plan that attacks the candidate's likely correctness and performance failure modes.

For engineering VAL-* or legacy engineering targets, do not use optimization outcome semantics. Pass only when the assigned benchmark, performance, correctness, and guardrail behavior required by the assignment and contract is proven with fresh evidence.

Inputs

Read:

  • Validation assignment.
  • The single assigned benchmark-related contract.
  • AGENTS.md.
  • Experiment ledger path cited by the assignment or contract.
  • Candidate artifact/ref/patch/checksum cited by the assignment, contract, or ledger.
  • Measurement protocol, correctness/guardrail commands, protected files, and source/baseline refs cited by the contract.
  • For optimization targets, also inspect visible runtime resources read-only when useful: changed files, benchmark scripts, fixtures, public outputs, baseline/reference artifacts, verifier-adjacent code, logs, and generated artifacts. Do not mutate these resources, copy protected artifacts into a submission, or make submitted code depend on evaluator-only locations.

If the assignment targets more than one benchmark-related assertion, fail the assignment as too broad and request attention.

Procedure

  1. Identify artifacts
    • Parent/baseline ref.
    • Candidate ref or patch artifact.
    • Benchmark command, run count, aggregation, variance policy.
    • Correctness and guardrail commands.
    • Protected benchmark/scoring/data/verifier files.

Read the full file on GitHub · 153 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 153 lines · 44 tokens per session scan A bf618457802d

Subscribe to this mod's changes

benchmark-validator is a skill published in the GitHub repository Intelligent-Internet/zenith (283 stars, last pushed 26d ago), licensed Apache-2.0. It adds 44 tokens to every session and 1,454 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

agent-harness-fault-injection

Use when an agent workflow needs deterministic recovery evidence for sandbox, MCP/tool, worker, checkpoint, memory, or orchestration failures.

sickn33/agentic-awesome-skills · 34 tokens

terraform-skill

Use when working with Terraform or OpenTofu - creating modules, writing tests (native test framework, Terratest), setting up CI/CD pipelines, reviewing configurations, choosing between testing approaches, debugging state issues, implementing security scanning (trivy, checkov), or making infrastructure-as-code…

agentscope-ai/QwenPaw · 62 tokens

xlsx

当电子表格文件是主要输入或输出时使用此技能。这意味着用户想要:打开、读取、编辑或修复现有的 .xlsx、.xlsm、.csv 或 .tsv 文件(例如添加列、计算公式、格式化、制图、清理混乱数据);从头创建新的电子表格或从其他数据源创建;或在表格文件格式之间进行转换。当用户通过名称或路径引用电子表格文件时特别触发——即使是随意提及(如"我下载目录里的 xlsx")——并且想对其进行操作或从中生成内容。也适用于将混乱的表格数据文件(格式错误的行、错位的表头、垃圾数据)清理或重构为规范的电子表格。交付物必须是电子表格文件。当主要交付物是 Word 文档、HTML 报告、独立 Python 脚本、数据库管道或 Google Sheets…

agentscope-ai/QwenPaw · 232 tokens

terraform-cli-setup

Terraform CLI 安装与初始化技能。当用户本地未安装 Terraform 时自动完成安装,确保 terraform 命令可用并能执行 init/validate。不负责 Provider 凭证配置,凭证在实际使用时由 terraform-skill 引导。.

agentscope-ai/QwenPaw · 58 tokens

pr-integration-test

Design, implement, and validate Intelligent Terminal integration tests for a target pull request or regression. Use when asked to add PR integration tests, convert a bug fix into E2E coverage, prove existing behavior still works, map tests to the release checklist, or verify E2E reports mark checklist cases complete.

microsoft/intelligent-terminal · 66 tokens

oma-scholar

Scholarly research companion using Knows sidecar spec (.knows.yaml). Generates, validates, reviews, queries, and compares structured research-paper sidecars, and fetches them from knows.academy. Use for academic literature search, survey synthesis, paper authoring assistance, and peer review with token-efficient…

first-fluke/oh-my-agent · 73 tokens