Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add ntorga/agent-starter-kit --skill coder-self-reviewgit clone --depth 1 https://github.com/ntorga/agent-starter-kitWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/ntorga/agent-starter-kit/coder-self-review)<a href="https://agentmods.dev/skills/ntorga/agent-starter-kit/coder-self-review"><img src="https://agentmods.dev/badge/skills/ntorga/agent-starter-kit/coder-self-review/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/ntorga/agent-starter-kit/coder-self-review"><img src="https://agentmods.dev/badge/skills/ntorga/agent-starter-kit/coder-self-review.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00024 | $0.02199 |
| Opus 5 | $0.00012 | $0.01099 |
| Sonnet 5 | $0.00005 | $0.00440 |
| Haiku 4.5 | $0.00002 | $0.00220 |
Grade A, and why
coder-self-review scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 110 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Purpose
Before delivering a handoff, the Coder evaluates its own output against the GRASP rubric. Each letter is scored 0, 1, or 2 with evidence quoted from the rubric and cited from actual work. The total determines whether to deliver, rewrite, or abort.
Procedure
-
Gather evidence. Honest self-review makes verification efficient. Accurate scorecards confirm fast; dishonest ones fail and re-run, wasting time and compute. Before scoring, run verification commands to gather proof. The examples below show common patterns — choose what provides the best evidence for your specific work.
Examples:
- Incomplete markers:
rg -n 'TODO|FIXME|HACK|stub|placeholder' <changed-files>— zero results supports higher scores - Test suite: run the project's test command — all passing supports higher scores
- Lint suppressions:
rg -n 'nolint|eslint-disable|@ts-ignore|# noqa|disable-line' <changed-files>— zero results or justified suppressions support higher scores - Secrets:
rg -n 'password|secret|api_key|token|PRIVATE.KEY' <changed-files>— zero results in source supports higher scores - Attack surface: identify every endpoint/handler accepting external input — verify each has explicit auth and input sanitization
- Run the project's linter if one exists.
These are examples, not mandates. Choose commands that provide the strongest proof for your work.
- Incomplete markers:
-
Score each criterion. Read the GRASP rubric below. For each letter, assign a score of 0, 1, or 2. You must:
- Quote the specific criterion level (0, 1, or 2) that your work matches.
- Cite evidence from your actual work — file paths read, commands run, test output, patterns checked. Generic claims like "I followed all guidelines" are not evidence and score 0.
-
Output the Scorecard. Fill in the scorecard below. This is not internal reasoning — this is your deliverable checkpoint.
-
Apply the hard-fail rule. If any letter scores 0, do not deliver — go to step 5 immediately.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 110 lines · 24 tokens per session scan A 26e830f31835
coder-self-review is a skill published in the GitHub repository ntorga/agent-starter-kit (142 stars, last pushed 2d ago), licensed MIT. It adds 24 tokens to every session and 2,199 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-13.
Other skills, from other repositories
python-testing
Guidelines for writing and running tests in the Agent Framework Python codebase. Use this when creating, modifying, or running tests.
foundry-hosted-agent-validation
Step-by-step process for validating a Python Foundry hosted agent sample (under python/samples/04-hosting/foundry-hosted-agents/) end to end — running it locally (native runtime and azd ai agent run) and after deploying it to an Azure AI Foundry project with azd. Use this when asked to validate a hosted agent sample.
verify-samples-tool
How to use the verify-samples tool to run, verify, and manage sample definitions in the Agent Framework repository. Use this when adding, updating, or running sample verification.
build-and-test
How to build and test .NET projects in the Agent Framework repository. Use this when verifying or testing changes.
verify-dotnet-samples
How to build, run and verify the .NET sample projects in the Agent Framework repository. Use this when a user wants to verify that the samples still function as expected.
regex-tester
Validate, test, and debug regular expressions by executing them against sample inputs. Use when asked to build, verify, or explain a regex pattern.