Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add AIDevGTM/gtm-cofounder --skill 17-review-the-workgit clone --depth 1 https://github.com/AIDevGTM/gtm-cofounderWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/aidevgtm/gtm-cofounder/17-review-the-work)<a href="https://agentmods.dev/skills/aidevgtm/gtm-cofounder/17-review-the-work"><img src="https://agentmods.dev/badge/skills/aidevgtm/gtm-cofounder/17-review-the-work/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/aidevgtm/gtm-cofounder/17-review-the-work"><img src="https://agentmods.dev/badge/skills/aidevgtm/gtm-cofounder/17-review-the-work.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00083 | $0.01173 |
| Opus 5 | $0.00042 | $0.00587 |
| Sonnet 5 | $0.00017 | $0.00235 |
| Haiku 4.5 | $0.00008 | $0.00117 |
Grade A, and why
review-the-work scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 13d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 57 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Review the work (the self-check gate)
The person who wrote it is the worst judge of whether it is good. Before this reaches the founder, stop being the author and become the skeptic.
Use this when: you are about to present any substantive deliverable, or the founder asks "is this actually good?" It is a gate, not a stage. It has no place in the linear sequence. Run it around any piece of work, every time.
The core idea
An agent that just wrote something is biased to ship it. That bias is how generic, plausible, quietly wrong work reaches a founder who does not yet know enough to catch it. So separate the two jobs: the author drafts, a different lens judges. You do not present work because you made it. You present it because it survived a skeptic.
Steal the one rule that makes this work: the author may submit the work, the author may not issue the verdict. Switch roles on purpose. Become the developer who is skeptical, the buyer who is busy, the reviewer who has seen a hundred of these, and try to break it before the market does.
How to run this (for the agent)
- Run it silently, before presenting. The founder should see the verdict and the work, not the whole audit.
- Give a clear verdict: PASS, or REVISE with the exact checks that failed and the specific fix for each. Never a vague "looks good."
- Do not rubber-stamp. If you cannot find a single weakness, you are not reading as the skeptic. Name the weakest point even in work you pass, so the founder knows where it is thin.
- When it fails, fix it and re-run the gate, then present. Do not hand the founder a list of problems you could have fixed yourself.
The standard (applies to everything)
- Specific, not puffery. No "powerful," "seamless," "best-in-class," "platform." Every claim carries a number, a name, or a proof, or it gets cut.
- The developer is the hero, not the product. If the founder's tool is the hero of the sentence, rewrite it.
- It speaks to the real ICP from the brief, not to "developers." A message for everyone lands on no one.
- It rests on validated facts. Anything load-bearing that is still
[assumption]is flagged out loud, not smuggled in as truth. - A skeptical developer could not immediately prove it false. Developers test claims and find the truth. If a dev would find the failing case in thirty seconds, you have not passed.
- Human voice. Reads like a person wrote it. No em-dashes, no AI tells.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 13d ago First seen · 57 lines · 83 tokens per session scan A dedcc811eaba
review-the-work is a skill published in the GitHub repository AIDevGTM/gtm-cofounder (276 stars, last pushed 5d ago), licensed MIT. It adds 83 tokens to every session and 1,173 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
git-phase-restore
Autonomous Git-based project phase restoration. Uses Git history (commits, tags, branches, diffs, semantic messages) to identify and restore any development phase automatically. Trigger for: "restore to when X worked", "go back to before Y broke", "show project phases", "undo the last feature", "roll back to [phase]"…
github-actions-trigger
Trigger GitHub Actions workflows from Claude Code. Auto-run tests, deploy, or validate when skills like run-pre-commit-checks complete.
e2b-sandboxes
E2B open-source cloud sandboxes for executing AI-generated code securely. USE FOR: run code in sandbox, e2b, code interpreter, execute AI code, secure code execution, isolated environment, AI coding agent execution, run untrusted code, cloud sandbox, code execution API, python sandbox cloud, AI agent code runner, safe…
screenshot-automation
Automatically captures, crops, and beautifies screenshots from URLs for your portfolio or docs.
exploratory-data-analysis
Perform bounded, local exploratory analysis of explicitly supported scientific files. Use for redacted CSV/TSV/JSON profiles; optional NumPy, HDF5, FASTA/FASTQ, and basic image metadata inspection; missingness/leakage audits; outlier and transformation sensitivity; and rigorous EDA report scaffolds. Other domain…
google-ads-audit
Google Ads account audit and business context setup. Use for account-health audits and business-context setup. Trigger on "audit my ads", "ads audit", "set up my ads", "onboard", "account overview", "how's my account", "ads health check", "what should I fix in my ads", or when the user is new to NotFair and hasn't run…