Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add jdrhyne/agent-skills --skill autoresearch-loopgit clone --depth 1 https://github.com/jdrhyne/agent-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/jdrhyne/agent-skills/autoresearch-loop)<a href="https://agentmods.dev/skills/jdrhyne/agent-skills/autoresearch-loop"><img src="https://agentmods.dev/badge/skills/jdrhyne/agent-skills/autoresearch-loop/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/jdrhyne/agent-skills/autoresearch-loop"><img src="https://agentmods.dev/badge/skills/jdrhyne/agent-skills/autoresearch-loop.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00120 | $0.01857 |
| Opus 5 | $0.00060 | $0.00928 |
| Sonnet 5 | $0.00024 | $0.00371 |
| Haiku 4.5 | $0.00012 | $0.00186 |
Grade A, and why
autoresearch-loop scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 61 lines — stays where its author put it; the contents beside it link to each section on GitHub.
autoresearch-loop
Generalize Karpathy's autoresearch into a domain-adaptive improvement loop. The agent discovers what to measure, then runs a disciplined propose → trial → keep-or-revert loop, maintaining an explicit ledger of what was tried, kept, discarded, and implemented.
Read DESIGN.md once at the start of a run for the full architecture and the domain-specific tensions (metric latency/noise/cost, Goodhart gaming, cost-per-trial, reversibility). The phases below are the operating procedure.
Runtime: the loop's mechanics (run a trial, parse the metric, score confidence, keep/commit or discard/revert) are handled by the arl CLI over a .auto/ session folder — a Claude-native port of pi-autoresearch's tools. Read references/runtime-contract.md for the .auto/ layout, the METRIC name=value contract, MAD confidence scoring, and the arl init|run|log|status commands. Invoke it as node scripts/arl.mjs <cmd> (or arl if on PATH).
When NOT to run
- There is no metric that can be measured repeatably and cheaply enough within a trial budget. A loop with no trustworthy metric chases noise — stop and say so.
- The change surface is irreversible or unsafe to mutate experimentally (production data, customer-facing irreversible actions) without an explicit revert procedure in the adapter.
Phase 0 — FRAME (metric discovery)
Given {project, goal, context} — follow the procedure in references/metric-discovery.md (restate the goal as an outcome → enumerate candidates on the proxy→outcome spectrum → score on six axes → choose primary + guardrails + strategy → red-team for gaming). In brief:
- Identify or select a domain adapter (
adapters/*.md). If none fits, draft an inline adapter followingreferences/domain-adapter-contract.md. - Propose candidate metrics (the adapter's menu is a prior, not the answer — reason from the goal). Score each on measurability, latency, noise, alignment, gameability, and cost. Prefer alignment over convenience for the primary.
- Pick ONE primary objective, a set of guardrail metrics that must not regress, a trial budget (wall-clock and/or cost per trial), and a stop condition (budget exhausted, plateau over K trials, or target hit).
- Choose the accept/reject strategy for this domain:
deterministic-delta(fast, low-noise),significance-test, orbandit(noisy/delayed/expensive — e.g. ads). - Name the Goodhart guards (guardrail metrics, holdout, periodic critic).
- Write
runs/<id>/CHARTER.md. Confirm it with the user before spending real budget if trials cost money or touch production.
What ships with it
22 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- adapters/bug-finding.md 3.4 KB
- adapters/code-generation.md 2.8 KB
- adapters/code-perf-audit.md 5.8 KB
- adapters/google-ads.md 3.2 KB
- adapters/web-onboarding.md 5.1 KB
- DESIGN.md 11 KB
- examples/code-perf-demo/.auto/checks.sh 110 B runs code
- examples/code-perf-demo/.auto/ideas.md 275 B
- examples/code-perf-demo/.auto/log.jsonl 1011 B
- examples/code-perf-demo/.auto/measure.sh 162 B runs code
- examples/code-perf-demo/.auto/prompt.md 634 B
- examples/code-perf-demo/bench.js 783 B runs code
- examples/code-perf-demo/gen.js 849 B runs code
- examples/code-perf-demo/RESULTS.md 682 B
- examples/code-perf-demo/solution.js 967 B runs code
- examples/code-perf-demo/test.js 1.2 KB runs code
- README.md 3.4 KB
- references/domain-adapter-contract.md 2.2 KB
- references/journal-schema.md 1.4 KB
- references/metric-discovery.md 4.2 KB
- references/runtime-contract.md 3.7 KB
- scripts/arl.mjs 14 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 61 lines · 120 tokens per session scan A 721e9fd122f6
autoresearch-loop is a skill published in the GitHub repository jdrhyne/agent-skills (241 stars, last pushed 10d ago), licensed MIT. It adds 120 tokens to every session and 1,857 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
ios-simulator
Verify and debug native, React Native, Expo, or Flutter apps on an iOS Simulator with agent-device. Use when an agent needs to launch an app, inspect its live UI, tap, type, scroll, validate a code change, collect failure evidence, or reproduce a workflow on an iPhone or iPad Simulator.
autonomous-run
Prepare, start, inspect, resume, or stop a finite local overnight coding run after a human has accepted a Wayfinder terminal spec; coordinates a declared Claude/Codex maker and independent checker without pushing, merging, or writing to external systems.
auto-qa
QAMESH project QA mesh — plan, run, report, and publish deterministic QA evidence.
monitor-patterns
Monitor tool usage patterns and grep --line-buffered compatibility guide.
experiment
Experiment loop for iterative metric-driven code optimization using XLOOP.
agent-presets
Domain-specific agent configurations for different project types.