eval-set

A project-specific test set for checking whether a retrieval system brings back the rules needed for different prompts. Retrieval means finding relevant guidance from a larger collection when it is needed.

In plain words
What is it for?
Use it after moving instructions from an always-loaded prompt into Clawness rules to compare results before and after the change using ranking and hit-rate scores.
Why use it?
It replaces visual guesswork with measurements that reveal when editing a prompt or rule makes important guidance harder to find.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/fullymiddleaged/clawness/eval-set
Any agent
npx skills add fullymiddleaged/Clawness --skill eval-set
Clone the repo
git clone --depth 1 https://github.com/fullymiddleaged/Clawness

Made for: Claude Code, Codex.

Per session 98 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,568 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00098 $0.01568
Opus 5 $0.00049 $0.00784
Sonnet 5 $0.00020 $0.00314
Haiku 4.5 $0.00010 $0.00157

Measured 2d ago against content hash 39523a8d4d47, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

eval-set scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/eval-set/SKILL.md · 131 lines

How it starts

The opening of the file, as written. The whole thing — 131 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Score your retrieval, don't eyeball it

When you move guidance out of an always-loaded base prompt (a bloated CLAUDE.md, an OpenClaw SOUL.md/AGENTS.md) into Clawness's ranked retrieval, you trade a guarantee for a probability: the content used to be present on every turn; now it surfaces only when the prompt is relevant enough to rank it. That trade is usually right — it is the whole point of /clawness:claude-md and /clawness:openclaw-audit — but it is only safe if you can check that the content still surfaces for the prompts that need it.

This skill builds that check. It is the same machinery Clawness gates its own corpus with: a labelled set of prompt → expected rule ID(s) cases, scored by MRR@k and hit-rate via clawness eval. The output is a number that moves when retrieval regresses, so a rule edit that quietly stops surfacing shows up instead of hiding until someone hits it in anger.

This is harness-agnostic — it evaluates Clawness retrieval, which is identical under Claude Code and under the OpenClaw adapter. Nothing here touches OpenClaw's own prompt.

When to run it

  • After trimming a base prompt into .clawness/rules/ — write a case for each thing the moved content used to guarantee, then confirm hit-rate is 1.0 before you delete the original. This is the verification step openclaw-audit/claude-md point at.
  • After editing or adding rules — re-run to confirm you didn't push an existing rule out of the top-k for prompts that depend on it.
  • As a CI gate, once the set is stable — the same --floor-mrr/--floor-hit floors Clawness uses on its own eval.

Steps

1. Create the case file

Author .clawness/eval/cases.json in the shape below (a filled-in copy of this skill's cases.template.json). Write it directly — the plugin root isn't reachable from skill Bash, so don't try to cp the template from the plugin dir.

The shape (identical to tests/ground_truth.json):

{
  "queries": [
    { "q": "how should I handle errors in our service layer", "expect": ["SVC-ERR-001"] }
  ]
}

Read the full file on GitHub · 131 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 131 lines · 98 tokens per session scan A 39523a8d4d47

Subscribe to this mod's changes

eval-set is a skill published in the GitHub repository fullymiddleaged/Clawness (3 stars, last pushed 4d ago), licensed MIT. It adds 98 tokens to every session and 1,568 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.