eval

A test runner for checking the project's stored fitness tests against its current state. Each probe reports PASS, REGRESSION, or SKIPPED, showing whether a previously fixed behavior still works.

In plain words
What is it for?
Use it to run all probes, one probe, or a selected test tier and update the project's RESULTS.md scoreboard.
Why use it?
It makes recurring problems visible instead of relying on manual checks. A regression identifies a previously passing test that has started to fail.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/mifunedev/openharness/eval
Any agent
npx skills add mifunedev/openharness --skill eval
Clone the repo
git clone --depth 1 https://github.com/mifunedev/openharness

Made for: Claude Code, Codex.

Per session 128 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 910 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00128 $0.00910
Opus 5 $0.00064 $0.00455
Sonnet 5 $0.00026 $0.00182
Haiku 4.5 $0.00013 $0.00091

Measured 2d ago against content hash 5f0fe00474ba, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

The scan reads SKILL.md. This mod also ships 1 executable file (run.sh), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.oh/skills/eval/SKILL.md · 63 lines

How it starts

The opening of the file, as written. The whole thing — 63 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Eval

The runner for the harness fitness function. It discovers .oh/evals/probes/*.sh, runs each against real state, and writes the .oh/evals/RESULTS.md scoreboard. A rectification is provably "done" when its probe is green; a recurrence shows up as a REGRESSION (was-PASS, now-fail) naming the # source: lesson. The full contract — 3-state exit oracle, header convention, correction-surface triage — is in .oh/evals/README.md.

Usage

bash .claude/skills/eval/run.sh                 # run the whole suite, rewrite RESULTS.md
bash .claude/skills/eval/run.sh --probe <id>    # run one probe, update only its row
bash .claude/skills/eval/run.sh --tier A        # run only Tier-A probes

Exit-code oracle (per probe): 0=PASS, 1=REGRESSION, 2=SKIPPED (not applicable — excluded from pass-rate), 124=TIMEOUT, other=ERROR. Each probe is wrapped in timeout 30s. Runner aggregate exit (the process $? of run.sh itself): 0 when no new green→red regression occurred this run, 1 when one or more new regressions were detected (${#regressions[@]} > 0). When invoked via the Bash tool as bash .claude/skills/eval/run.sh, the agent caller reads $? directly to gate on success — the printed REGRESSIONS (...) stdout block and per-probe stderr lines remain the human-readable signal. Note: the eval-weekly cron is an intentional legacy caller that appends || true then greps stdout; it does not consume the exit code by design — this is not a bug.

What the runner does

  1. Discover + run every probe matching the filters; extract # tier: / # source: via the exact header grep.
  2. Compute the delta vs the prior RESULTS.md row. First run (no prior row) emits new-pass/new-fail and raises NO regression without prior state.
  3. Surface regressions — any PASS → (REGRESSION|TIMEOUT|ERROR) transition is printed first, naming the probe's source.
  4. Rewrite RESULTS.md atomically — build the full scoreboard into a temp sibling file (RESULTS.md.tmp.$$) and replace the live file in one mv -f (never truncate-then-append in place), so a crash or concurrent run can't leave a partial scoreboard. Overwrite the row for each probe run; carry prior rows for probes not run this invocation from a pre-write snapshot (RESULTS_ORIG) captured before the rewrite — not the live file — so a filtered run never erases untouched rows and the scoreboard stays complete.

Read the full file on GitHub · 63 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 63 lines · 128 tokens per session scan A 5f0fe00474ba

Subscribe to this mod's changes

eval is a skill published in the GitHub repository mifunedev/openharness (36 stars, last pushed 2d ago), licensed Apache-2.0. It adds 128 tokens to every session and 910 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

data-visualization

Use for creating publication-quality charts and multi-panel analysis summaries. Triggers when tasks involve visualizing data, plotting results, creating charts, or producing visual reports from analysis output.

langchain-ai/deepagents · 40 tokens

cuml-machine-learning

Use for GPU-accelerated machine learning on tabular data using NVIDIA cuML. Triggers when tasks involve classification, regression, clustering, dimensionality reduction, or model training on datasets.

langchain-ai/deepagents · 43 tokens

blog-post

Writes and structures long-form blog posts, creates tutorial outlines, and optimizes content for SEO with cover image generation. Use when the user asks to write a blog post, article, how-to guide, tutorial, technical writeup, thought leadership piece, or long-form content.

langchain-ai/deepagents · 58 tokens

social-media

Drafts engaging social media posts, writes hooks, suggests hashtags, creates thread structures, and generates companion images. Use when the user asks to write a LinkedIn post, tweet, Twitter/X thread, social media caption, social post, or repurpose content for social platforms.

langchain-ai/deepagents · 58 tokens

remember

Review the current conversation and capture valuable knowledge — best practices, coding conventions, architecture decisions, workflows, and user feedback — into persistent memory (AGENTS.md) or reusable skills. Use when the user says: (1) remember this, (2) save what we learned, (3) update memory, (4) capture…

langchain-ai/deepagents · 71 tokens

textual-screenshot

Capture a Textual terminal UI as an SVG using its headless test harness. Use when asked to make, attach, or preview a screenshot of deepagents-code/dcode or another Textual app, visually verify a TUI state, or render a modal, screen, or widget without a desktop or browser.

langchain-ai/deepagents · 67 tokens