Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/fullymiddleaged/clawness/eval-setnpx skills add fullymiddleaged/Clawness --skill eval-setgit clone --depth 1 https://github.com/fullymiddleaged/ClawnessWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00098 | $0.01568 |
| Opus 5 | $0.00049 | $0.00784 |
| Sonnet 5 | $0.00020 | $0.00314 |
| Haiku 4.5 | $0.00010 | $0.00157 |
Grade A, and why
eval-set scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 131 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Score your retrieval, don't eyeball it
When you move guidance out of an always-loaded base prompt (a bloated CLAUDE.md, an
OpenClaw SOUL.md/AGENTS.md) into Clawness's ranked retrieval, you trade a guarantee
for a probability: the content used to be present on every turn; now it surfaces only
when the prompt is relevant enough to rank it. That trade is usually right — it is the
whole point of /clawness:claude-md and /clawness:openclaw-audit — but it is only safe
if you can check that the content still surfaces for the prompts that need it.
This skill builds that check. It is the same machinery Clawness gates its own corpus with:
a labelled set of prompt → expected rule ID(s) cases, scored by MRR@k and hit-rate via
clawness eval. The output is a number that moves when retrieval regresses, so a rule edit
that quietly stops surfacing shows up instead of hiding until someone hits it in anger.
This is harness-agnostic — it evaluates Clawness retrieval, which is identical under Claude Code and under the OpenClaw adapter. Nothing here touches OpenClaw's own prompt.
When to run it
- After trimming a base prompt into
.clawness/rules/— write a case for each thing the moved content used to guarantee, then confirm hit-rate is 1.0 before you delete the original. This is the verification stepopenclaw-audit/claude-mdpoint at. - After editing or adding rules — re-run to confirm you didn't push an existing rule out of the top-k for prompts that depend on it.
- As a CI gate, once the set is stable — the same
--floor-mrr/--floor-hitfloors Clawness uses on its own eval.
Steps
1. Create the case file
Author .clawness/eval/cases.json in the shape below (a filled-in copy of this skill's
cases.template.json). Write it directly — the plugin root isn't reachable from skill
Bash, so don't try to cp the template from the plugin dir.
The shape (identical to tests/ground_truth.json):
{
"queries": [
{ "q": "how should I handle errors in our service layer", "expect": ["SVC-ERR-001"] }
]
}
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 131 lines · 98 tokens per session scan A 39523a8d4d47
eval-set is a skill published in the GitHub repository fullymiddleaged/Clawness (3 stars, last pushed 4d ago), licensed MIT. It adds 98 tokens to every session and 1,568 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
issue-triage
Issue triage: audit open issues, categorize, detect duplicates, cross-ref PRs, risk assessment, post comments. Args: "all" for deep analysis of all, issue numbers to focus (e.g. "42 57"), "en"/"fr" for language, no arg = audit only in French.
ship
Build, commit, push & version bump workflow - automates the complete release cycle.
pr-review
Batch review des PRs RTK par ordre de complexité croissante (XS → S → M → L). Pour chaque PR : vérifie l'état (conflits, CLA, reviews), lit le diff complet, analyse le code en contexte, présente un résumé avec lien + taille + recommandation. Attend validation explicite avant tout merge. Poste des commentaires…
rtk-triage
Triage complet RTK : exécute issue-triage + pr-triage en parallèle, puis croise les données pour détecter doubles couvertures, trous sécurité, P0 sans PR, et conflits internes. Sauvegarde dans claudedocs/RTK-YYYY-MM-DD.md. Args: "en"/"fr" pour la langue (défaut: fr), "save" pour forcer la sauvegarde.
code-simplifier
Review RTK Rust code for idiomatic simplification. Detects over-engineering, unnecessary allocations, verbose patterns. Applies Rust idioms without changing behavior.
security-guardian
CLI security expert for RTK - command injection, shell escaping, hook security.