Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add aaronnat23/disp8ch --skill experiment-loopgit clone --depth 1 https://github.com/aaronnat23/disp8chWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/aaronnat23/disp8ch/experiment-loop)<a href="https://agentmods.dev/skills/aaronnat23/disp8ch/experiment-loop"><img src="https://agentmods.dev/badge/skills/aaronnat23/disp8ch/experiment-loop.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00000 | $0.00831 |
| Opus 5 | $0.00000 | $0.00415 |
| Sonnet 5 | $0.00000 | $0.00166 |
| Haiku 4.5 | $0.00000 | $0.00083 |
Grade A, and why
experiment-loop scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 66 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Experiment Loop (Metric-Driven Optimization)
Autonomous benchmark-driven experiment loop: propose a change, measure the metric, keep what improves it, revert what doesn't, repeat. Git is the ledger — every kept improvement is a commit. Every discarded attempt is a clean revert.
Setup
Before iterating, call init_experiment to configure the session:
metric_name: the scalar to optimize (e.g.test_duration_ms,accuracy_pct,bundle_size_kb)metric_direction:"minimize"or"maximize"objective: plain-English goal (e.g. "Reduce test suite wall-clock time below 10s")benchmark_command: shell command that must printMETRIC <name>=<number>to stdoutchecks_command(optional): correctness guard (e.g.npm testorpython -m pytest) — a failed check blocks a keep even if the metric improved (Goodhart's Law prevention)
Loop
- Propose: read
autoresearch.ideas.mdfor queued ideas; pick the most promising and describe the change - Implement: make the code change using
write_fileorbash_exec - Measure: call
run_experimentwith a description of what was tried - Decide: call
log_experimentwith:decision="keep"if metric improved AND checks passed → git commitdecision="discard"if metric regressed or was flat → git revertdecision="checks_failed"if metric improved but correctness check failed → revertdecision="crash"if the benchmark itself crashed → revert
- Record ideas: append promising-but-deferred ideas to
autoresearch.ideas.md - Repeat
Benchmark Contract
The benchmark command MUST print at least one METRIC name=number line to stdout:
METRIC test_duration_ms=4231
METRIC memory_mb=128
Secondary metrics are captured automatically. The primary metric is whatever metric_name was set to in init_experiment.
Segment-Aware Baselines
Call init_experiment again mid-session to start a new baseline segment — useful when pivoting to a different optimization goal. Old results are preserved in autoresearch.jsonl with their original segment index.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 66 lines · 0 tokens per session scan A 2f14aba39335
experiment-loop is a skill published in the GitHub repository aaronnat23/disp8ch (99 stars, last pushed 4d ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 831 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
sup-coding
Coding workflow recipe (brainstorm → design → spec → plan → implement → review → verify → finish).
setup-pre-commit
Set up Husky pre-commit hooks with lint-staged (Prettier), type checking, and tests in the current repo. Use when user wants to add pre-commit hooks, set up Husky, configure lint-staged, or add commit-time formatting/typechecking/testing.
workflow-patterns
Use this skill when implementing tasks according to Conductor's TDD workflow, handling phase checkpoints, managing git commits for tasks, or understanding the verification protocol.
meta-pre-commit-quality-gate
Run three quality gates (ruff + mypy + pytest) in parallel over the staged diff, then arbitrate a single BLOCK/APPROVE verdict. Use before committing changes locally when you want a comprehensive pre-commit gate beyond per-file linting — exactly the same gate set CI enforces.
commit
Create git commits with good messages. Use when user says "commit", "create commit", or asks to commit changes.
nw-quality-framework
Quality gates - 11 commit readiness gates, build/test protocol, validation checkpoints, and quality metrics.