Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/epicsagas/epic-harness/evalnpx skills add epicsagas/epic-harness --skill evalgit clone --depth 1 https://github.com/epicsagas/epic-harnessWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/epicsagas/epic-harness/eval)<a href="https://agentmods.dev/skills/epicsagas/epic-harness/eval"><img src="https://agentmods.dev/badge/skills/epicsagas/epic-harness/eval.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00041 | $0.02089 |
| Opus 5 | $0.00020 | $0.01045 |
| Sonnet 5 | $0.00008 | $0.00418 |
| Haiku 4.5 | $0.00004 | $0.00209 |
Grade A, and why
eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 210 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Eval — Quality & Regression Gate
CRITICAL: Run HARNESS_DIR=$(epic path) first. Never use .harness/ in the project directory.
When to Trigger
- Before
/shipcreates a PR (automatic if eval.yaml exists) - After
/gocompletes a feature - On explicit
/evalcommand - When user mentions "regression", "baseline", "eval suite", "quality check"
- CI:
make evalorepic eval --json
Execution Modes
4 dimensions run in parallel where possible:
- eval:correctness — Test pass rate, mutation score, assertion density
- eval:performance — Throughput, latency, memory (opt-in)
- eval:quality — Lint, code quality, LLM-as-judge
- eval:regression — Baseline comparison, score deltas
Process
Step 0: Prerequisites
HARNESS_DIR=$(epic path)
If $HARNESS_DIR/eval/eval.yaml does not exist, run scaffold:
epic eval --init
Read the config:
cat $HARNESS_DIR/eval/eval.yaml
Step 0.5: Scaffold benchmarks (when no benchmark infrastructure exists)
If eval.yaml has benchmarks: [] and no benchmark files are found in the project:
-
Generate stub files using the CLI:
epic eval --scaffoldSupported stacks (auto-detected from project markers):
Stack Detected by Generated file Output format Rust Cargo.tomlbenches/eval_harness.rscriterion (exit code) Python pyproject.toml/setup.pybenchmarks/eval_runner.pyJSON composite TypeScript tsconfig.jsonbenchmarks/eval.tsJSON composite Node.js package.jsonbenchmarks/eval.mjsJSON composite Go go.modbenchmarks/eval_test.goJSON composite Java pom.xml/build.gradlebenchmarks/EvalBenchmark.javaexit code Kotlin build.gradle.ktsbenchmarks/EvalBenchmark.ktexit code Ruby Gemfilebenchmarks/eval_benchmark.rbJSON composite PHP composer.jsonbenchmarks/eval_benchmark.phpJSON composite C# *.csproj/*.slnBenchmarks/EvalBenchmark.csJSON composite Swift Package.swiftbenchmarks/EvalBenchmark.swiftJSON composite Elixir mix.exsbenchmarks/eval_benchmark.exsJSON composite C++ CMakeLists.txtbenchmarks/eval_benchmark.cppexit code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 210 lines · 41 tokens per session scan A 8c0062e70d10
eval is a skill published in the GitHub repository epicsagas/epic-harness (17 stars, last pushed today), licensed Apache-2.0. It adds 41 tokens to every session and 2,089 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
mcporter
List, auth, and call MCP servers/tools from the terminal.
article-writing
Write articles, guides, blog posts, tutorials, newsletter issues, and other long-form content in a distinctive voice derived from supplied examples or brand guidance. Use when the user wants polished written content longer than a paragraph, especially when voice consistency, structure, and credibility matter.
mem0-oss-to-platform
Plan and then execute a migration of a project from the mem0 open-source / self-hosted SDK (the local Memory class) to the mem0 Platform / hosted / managed SDK (the MemoryClient class). Use this whenever a developer wants to move, switch, or migrate their mem0 usage off OSS/self-hosted to the hosted API — e.g.…
complete-partial-pr
Evaluate and complete an issue or PR where the submitted patch fixes only a narrow symptom of the reported pain point. Use when a contribution may miss adjacent integration surfaces, provider/spec semantics, roundtrip behavior, tests, docs, or historical maintainer decisions.
deploy-docker-compose
Run the Omnigent server as a Docker compose stack (server + Postgres) on any Docker host — your laptop, a VPS, EC2 by hand, or as the base layer of any container-platform deploy. Invoke when the user wants to build the image, bring up the compose stack, debug the stack on a host they already have, or extend the stack…
mapping-to-snomed
Maps clinical concept spans extracted by OpenMed to SNOMED CT concepts through a USER-SUPPLIED terminology server (the user's own Ontoserver, Snowstorm, or UMLS/UTS), never a bundled vocabulary. Use when the user wants to code findings, disorders, procedures, body structures, or substances to SNOMED CT, run an ECL…