Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/adrianco/retort/evaluate-runnpx skills add adrianco/retort --skill evaluate-rungit clone --depth 1 https://github.com/adrianco/retortWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/adrianco/retort/evaluate-run)<a href="https://agentmods.dev/skills/adrianco/retort/evaluate-run"><img src="https://agentmods.dev/badge/skills/adrianco/retort/evaluate-run.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00045 | $0.04380 |
| Opus 5 | $0.00023 | $0.02190 |
| Sonnet 5 | $0.00009 | $0.00876 |
| Haiku 4.5 | $0.00005 | $0.00438 |
Grade A, and why
evaluate-run scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 347 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Evaluate Retort Run
Overview
A retort run produces a workspace directory — generated source code for one factor-level combination — archived under <experiment>/runs/<cell>/rep<N>/. This skill evaluates that workspace against the task spec it was asked to implement, captures quantitative and qualitative findings, and writes results in a format comparable across runs.
This is the per-run counterpart to pourpoise's evaluate-attempt, adapted for retort's DoE structure: instead of ad-hoc attempts, each run is a point in a design matrix.
Parameters
- run_dir (required): Path to the archived run workspace, e.g.
experiment-1/runs/language=rust_model=opus_tooling=beads/rep2/ - output_file (optional, default:
{run_dir}/evaluation.md): Where to write the human-readable report - findings_file (optional, default:
{run_dir}/findings.jsonl): Where to write structured findings (one JSON object per line) suitable forfile-run-issues
Inputs You Can Rely On
Each run_dir is laid out by retort's LocalRunner and contains:
| File | Purpose |
|---|---|
TASK.md |
Task spec — the prompt the agent received. This is the "requirements" source of truth. |
stack.json |
{"language": ..., "agent": ..., "framework": ...} — the factor levels for this run |
| All generated source files | Exactly as the agent left them |
Possibly .beads/ |
Only if tooling=beads was in effect — the agent used bd for tracking |
| Possibly build artifacts | node_modules/, target/, __pycache__/, etc. |
The retort database (experiment-<N>/retort.db) also holds this run's ExperimentRun + RunResult rows. You MAY query it read-only for cross-checking scores; you MUST NOT write to it.
Steps
1. Verify the run workspace
test -d "{run_dir}" || { echo "run_dir missing"; exit 1; }
test -f "{run_dir}/TASK.md" || { echo "TASK.md missing — not a retort workspace"; exit 1; }
test -f "{run_dir}/stack.json" || { echo "stack.json missing"; exit 1; }
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 347 lines · 45 tokens per session scan A 5ad1ce901284
evaluate-run is a skill published in the GitHub repository adrianco/retort (199 stars, last pushed 4d ago), licensed Apache-2.0. It adds 45 tokens to every session and 4,380 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
brainstorming
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.
auto-perf-optimize
Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.
chat-perf
Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…