Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/sammcj/agentic-coding/iterative-refinementnpx skills add sammcj/agentic-coding --skill iterative-refinementgit clone --depth 1 https://github.com/sammcj/agentic-codingWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/sammcj/agentic-coding/iterative-refinement)<a href="https://agentmods.dev/skills/sammcj/agentic-coding/iterative-refinement"><img src="https://agentmods.dev/badge/skills/sammcj/agentic-coding/iterative-refinement.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00100 | $0.02795 |
| Opus 5 | $0.00050 | $0.01398 |
| Sonnet 5 | $0.00020 | $0.00559 |
| Haiku 4.5 | $0.00010 | $0.00280 |
Grade A, and why
iterative-refinement scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 127 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Iterative refinement with verifiable conditions
Improve any system (script, pipeline, prompt, doc, config, dataset) by looping against measurable pass/fail conditions while keeping the main thread's context lean. Task-agnostic. The aim is to make "is it good yet?" a single repeatable command, catch each regression at the edit that caused it, and keep load-bearing reasoning in cheap, auditable steps.
Use this as a toolkit, not a script. Each method below earns its place by what it prevents, and that reasoning is stated inline so you can judge when it applies. Reach for the methods the situation calls for, scale them to the stakes, and adapt or skip what doesn't fit; they compose well, but no fixed subset is mandatory and this skill can't anticipate every task you'll point it at. When a method clearly fits, lean into it fully rather than half-applying it. The judgement of which to use, and how hard, stays yours.
The loop
- Turn the goal into a pass/fail rubric with explicit thresholds. "Useful" isn't checkable; "contamination < 2%, both naive and effective ratios reported, every input row classified, prints
Overall: PASS" is. - Bake the rubric into the artifact as a self-check it prints. The artifact computes its own metrics and prints
[PASS]/[FAIL]per condition, so the check can't drift out of sync with the code the way an external checklist does. Now any party (you, a subagent, the user) re-verifies with one command. - Change one layer, re-run the check. Catch each regression at the edit that caused it, not three edits later.
- Loop on a fast slice, not the full dataset. Size the slice so a run takes seconds, not minutes. Run the full corpus only at checkpoints, when a layer is structurally complete and at sign-off, since that's where rare signals and period splits actually appear and the cost is justified.
- Keep a frozen held-out slice the loop never touches, and confirm against it before sign-off. Looping hard on one slice optimises the rubric for that slice (Goodhart on your own metric); the held-out slice is what proves the gain generalised rather than memorised.
- When all conditions pass but value remains, write a stricter rubric and loop again. Stop when a cycle yields nothing material.
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 127 lines · 100 tokens per session scan A b10046efe8fb
iterative-refinement is a skill published in the GitHub repository sammcj/agentic-coding (159 stars, last pushed yesterday), licensed Apache-2.0. It adds 100 tokens to every session and 2,795 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
strands-review
Local preview of the strands-agents/devtools /strands review agent. Body is the upstream Task Reviewer SOP verbatim — do not paraphrase. Use when the user types /strands-review, asks for a "strands review" of a PR, or wants to anticipate what the remote /strands review GitHub Action will flag. Findings are close but…
pr-writer
Generates pull request titles and descriptions. Use when the user asks to create, open, write, draft, or generate a PR, pull request, or merge request description.
docs-writer
Draft or rewrite Strands Agents documentation pages. Use when writing new doc pages, rewriting pages that failed audit, drafting sections for existing pages, or writing blog posts and release notes about Strands. Also triggers on "write a doc", "draft a page", "rewrite the quickstart", "add a tutorial for X"…
pre-push
Runs the local equivalent of the CI merge gate before you push. Detects which areas (Python, TypeScript, docs) your changes touch, auto-fixes what it can, then runs only those checks. Use when the user asks to run pre-push checks, get push-ready, verify changes before pushing or opening a PR, "make sure CI will pass"…
docs-planner
Identify documentation gaps and prioritize the docs backlog. Use when planning a docs improvement sprint, after signals surface repeated friction, when new SDK features ship without docs, or for periodic health assessment. Also triggers on "plan docs work", "what docs need writing", "prioritize the backlog", "docs…
pr-feedback
Fetches PR review feedback and inline comments, categorizes them, and presents options to the user. Use when the user asks to get, read, address, or fix review comments on a pull request.