Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/mifunedev/openharness/evalnpx skills add mifunedev/openharness --skill evalgit clone --depth 1 https://github.com/mifunedev/openharnessWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00128 | $0.00910 |
| Opus 5 | $0.00064 | $0.00455 |
| Sonnet 5 | $0.00026 | $0.00182 |
| Haiku 4.5 | $0.00013 | $0.00091 |
Grade A, and why
eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 63 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Eval
The runner for the harness fitness function. It discovers .oh/evals/probes/*.sh,
runs each against real state, and writes the .oh/evals/RESULTS.md scoreboard. A
rectification is provably "done" when its probe is green; a recurrence shows up as
a REGRESSION (was-PASS, now-fail) naming the # source: lesson. The full
contract — 3-state exit oracle, header convention, correction-surface triage — is
in .oh/evals/README.md.
Usage
bash .claude/skills/eval/run.sh # run the whole suite, rewrite RESULTS.md
bash .claude/skills/eval/run.sh --probe <id> # run one probe, update only its row
bash .claude/skills/eval/run.sh --tier A # run only Tier-A probes
Exit-code oracle (per probe): 0=PASS, 1=REGRESSION, 2=SKIPPED (not
applicable — excluded from pass-rate), 124=TIMEOUT, other=ERROR. Each probe is
wrapped in timeout 30s. Runner aggregate exit (the process $? of run.sh
itself): 0 when no new green→red regression occurred this run, 1 when one or
more new regressions were detected (${#regressions[@]} > 0). When invoked via the
Bash tool as bash .claude/skills/eval/run.sh, the agent caller reads $? directly
to gate on success — the printed REGRESSIONS (...) stdout block and per-probe stderr
lines remain the human-readable signal. Note: the eval-weekly cron is an intentional
legacy caller that appends || true then greps stdout; it does not consume the exit
code by design — this is not a bug.
What the runner does
- Discover + run every probe matching the filters; extract
# tier:/# source:via the exact header grep. - Compute the delta vs the prior
RESULTS.mdrow. First run (no prior row) emitsnew-pass/new-failand raises NO regression without prior state. - Surface regressions — any
PASS → (REGRESSION|TIMEOUT|ERROR)transition is printed first, naming the probe'ssource. - Rewrite
RESULTS.mdatomically — build the full scoreboard into a temp sibling file (RESULTS.md.tmp.$$) and replace the live file in onemv -f(never truncate-then-append in place), so a crash or concurrent run can't leave a partial scoreboard. Overwrite the row for each probe run; carry prior rows for probes not run this invocation from a pre-write snapshot (RESULTS_ORIG) captured before the rewrite — not the live file — so a filtered run never erases untouched rows and the scoreboard stays complete.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 63 lines · 128 tokens per session scan A 5f0fe00474ba
eval is a skill published in the GitHub repository mifunedev/openharness (36 stars, last pushed 2d ago), licensed Apache-2.0. It adds 128 tokens to every session and 910 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
data-visualization
Use for creating publication-quality charts and multi-panel analysis summaries. Triggers when tasks involve visualizing data, plotting results, creating charts, or producing visual reports from analysis output.
cuml-machine-learning
Use for GPU-accelerated machine learning on tabular data using NVIDIA cuML. Triggers when tasks involve classification, regression, clustering, dimensionality reduction, or model training on datasets.
blog-post
Writes and structures long-form blog posts, creates tutorial outlines, and optimizes content for SEO with cover image generation. Use when the user asks to write a blog post, article, how-to guide, tutorial, technical writeup, thought leadership piece, or long-form content.
social-media
Drafts engaging social media posts, writes hooks, suggests hashtags, creates thread structures, and generates companion images. Use when the user asks to write a LinkedIn post, tweet, Twitter/X thread, social media caption, social post, or repurpose content for social platforms.
remember
Review the current conversation and capture valuable knowledge — best practices, coding conventions, architecture decisions, workflows, and user feedback — into persistent memory (AGENTS.md) or reusable skills. Use when the user says: (1) remember this, (2) save what we learned, (3) update memory, (4) capture…
textual-screenshot
Capture a Textual terminal UI as an SVG using its headless test harness. Use when asked to make, attach, or preview a screenshot of deepagents-code/dcode or another Textual app, visually verify a TUI state, or render a modal, screen, or widget without a desktop or browser.