benchmark

A review gate that decides whether one completed change improved or preserved a project’s tested capabilities without lowering existing checks, and whether its complexity is justified.

In plain words
What is it for?
Use it after a change to run the project’s regression checks and capability benchmark, then receive a BENEFICIAL or NOT-BENEFICIAL verdict.
Why use it?
It separates useful progress from changes that only add testing machinery. It produces one clear result instead of leaving the value of the change ambiguous.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/mifunedev/openharness/benchmark
Any agent
npx skills add mifunedev/openharness --skill benchmark
Clone the repo
git clone --depth 1 https://github.com/mifunedev/openharness

Made for: Claude Code, Codex.

Per session 212 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,030 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00212 $0.02030
Opus 5 $0.00106 $0.01015
Sonnet 5 $0.00042 $0.00406
Haiku 4.5 $0.00021 $0.00203

Measured 3d ago against content hash 572bcfc7a05e, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

benchmark scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.oh/skills/benchmark/SKILL.md · 165 lines

How it starts

The opening of the file, as written. The whole thing — 165 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Benchmark — progress-ceiling verdict gate

The benchmark gate, part of /spec execute's improve tail in .oh/skills/spec/SKILL.md. It answers one question: was this change actually beneficial — did it move or hold the capability ceiling without breaking the regression floor, and is it worth its complexity? — and emits exactly one verdict.

Core principle: compose, don't re-derive — and judge OUTCOMES, not machinery. This skill owns the verdict, not the instruments. The regression floor is /eval; the progress ceiling is the capability benchmark (.oh/evals/capability/). /benchmark runs both and integrates them into a single BENEFICIAL / NOT-BENEFICIAL. Adding machinery is not progress — a change that grows the harness but does not move the capability benchmark is NOT-BENEFICIAL by definition.

Not /audit implementation. /audit implementation is the per-unit floor gate (does this one impl satisfy its task graph and is it promotable?). /benchmark is the ceiling gate (did the harness get better?). Distinct instruments, distinct question — see .oh/evals/capability/README.md § Ceiling vs. floor. /benchmark consults /eval; it does not replace or fork it.


Inputs

Arg Meaning
--base <ref> The counterfactual to score against — the state without this change. Defaults to the merge-base with development (i.e. "the repo before this change").
--cycles <N> Window for the redirect signal (§ Redirect signal). Default 3: a benchmark flat for N consecutive cycles while machinery grows trips the human-redirect flag.

The two signals (fail-fast, in order)

Run in order; the first signal that decides NOT-BENEFICIAL ends the gate. Only when both signals clear is the verdict BENEFICIAL. A signal that is missing or ambiguous is treated as NOT-BENEFICIAL, never as beneficial (honest exits).

Signal 1 — Regression floor (/eval)

A change that breaks the floor is never beneficial, whatever it claims to add. Gate on the runner's exit code + delta, not its prose — and read the cycle's single run rather than launching a third one:

Read the full file on GitHub · 165 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 165 lines · 212 tokens per session scan A 572bcfc7a05e

Subscribe to this mod's changes

benchmark is a skill published in the GitHub repository mifunedev/openharness (36 stars, last pushed 3d ago), licensed Apache-2.0. It adds 212 tokens to every session and 2,030 once invoked, about $0.0011 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

data-visualization

Use for creating publication-quality charts and multi-panel analysis summaries. Triggers when tasks involve visualizing data, plotting results, creating charts, or producing visual reports from analysis output.

langchain-ai/deepagents · 40 tokens

cuml-machine-learning

Use for GPU-accelerated machine learning on tabular data using NVIDIA cuML. Triggers when tasks involve classification, regression, clustering, dimensionality reduction, or model training on datasets.

langchain-ai/deepagents · 43 tokens

blog-post

Writes and structures long-form blog posts, creates tutorial outlines, and optimizes content for SEO with cover image generation. Use when the user asks to write a blog post, article, how-to guide, tutorial, technical writeup, thought leadership piece, or long-form content.

langchain-ai/deepagents · 58 tokens

social-media

Drafts engaging social media posts, writes hooks, suggests hashtags, creates thread structures, and generates companion images. Use when the user asks to write a LinkedIn post, tweet, Twitter/X thread, social media caption, social post, or repurpose content for social platforms.

langchain-ai/deepagents · 58 tokens

remember

Review the current conversation and capture valuable knowledge — best practices, coding conventions, architecture decisions, workflows, and user feedback — into persistent memory (AGENTS.md) or reusable skills. Use when the user says: (1) remember this, (2) save what we learned, (3) update memory, (4) capture…

langchain-ai/deepagents · 71 tokens

textual-screenshot

Capture a Textual terminal UI as an SVG using its headless test harness. Use when asked to make, attach, or preview a screenshot of deepagents-code/dcode or another Textual app, visually verify a TUI state, or render a modal, screen, or widget without a desktop or browser.

langchain-ai/deepagents · 67 tokens