Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add aneja5/forge-skills --skill writing-skillsgit clone --depth 1 https://github.com/aneja5/forge-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/aneja5/forge-skills/writing-skills)<a href="https://agentmods.dev/skills/aneja5/forge-skills/writing-skills"><img src="https://agentmods.dev/badge/skills/aneja5/forge-skills/writing-skills/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/aneja5/forge-skills/writing-skills"><img src="https://agentmods.dev/badge/skills/aneja5/forge-skills/writing-skills.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00053 | $0.02779 |
| Opus 5 | $0.00026 | $0.01389 |
| Sonnet 5 | $0.00011 | $0.00556 |
| Haiku 4.5 | $0.00005 | $0.00278 |
Grade A, and why
writing-skills scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 211 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Writing Skills
Overview
Writing a forge-skill IS test-driven development applied to documentation.
You write a pressure scenario (test case), watch a fresh agent fail at it without the skill (RED), write the skill (production code), watch the agent comply (GREEN), then close loopholes the agent invents (REFACTOR).
Adapted from Superpowers' writing-skills. See tests/METHODOLOGY.md for the full testing methodology.
When to Use
- Creating a new skill in
skills/<name>/SKILL.md - Editing an existing skill's body, frontmatter, or supporting files
- Promoting a one-off prompt that's been used 3+ times into a reusable skill
- Verifying a skill works before opening a PR
When NOT to Use
- Project-specific conventions — those go in CLAUDE.md, not a skill
- Pure reference material (templates, checklists) that doesn't enforce any rule — put it in
references/ - One-off solutions to one-off problems
The Iron Law
No skill ships without a failing test first.
Applies to new skills AND edits. If you can't show the baseline failure that motivated the skill, you don't know if the skill prevents the right failure.
Wrote the skill before running the test? Delete it. Start over. No exceptions:
- Not for "simple additions"
- Not for "just adding a section"
- Not for "documentation updates"
- Don't keep the un-tested draft as reference
Common Rationalizations
| Excuse | Reality |
|---|---|
| "It's just a documentation file, no need to test" | Untested skills have unstated assumptions. Test catches them. |
| "I'm just adding a rationalization to the table" | New rationalizations come from real session failures. If you didn't see one, don't add one. |
| "The skill is obviously clear to me" | Clear to you ≠ clear to a fresh agent. Test it. |
| "I'll test if a user reports a problem" | By then the skill has already shipped broken. Test first. |
| "Three pressures is too many for one scenario" | One pressure agents resist easily. Three pressures forces the genuine choice. |
| "RED passed but that means the skill works" | RED passing means the scenario isn't strong enough OR the skill prevents nothing useful. |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 211 lines · 53 tokens per session scan A db639282adcf
writing-skills is a skill published in the GitHub repository aneja5/forge-skills (3 stars, last pushed 3mo ago), licensed MIT. It adds 53 tokens to every session and 2,779 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
tdd
Use when implementing features or bug fixes test-first.
auto-loop
TDD-based autonomous development loop with checkpoint recovery and observability changelog.
workflow
Run the complete 5-step development workflow: focus problem → prevent over-development → test-first (TDD) → document → smart commit. Use when starting a new feature, or when the user runs /workflow or asks for the full development flow.
test-first
Drive one feature through a strict TDD Red-Green-Refactor cycle with checklists for each phase. Use when implementing new functionality test-first, or when the user runs /test-first.
test-audit
Audit test suites for T1-T4 violations using AST analysis, mock detection, and multi-stage synthesis. Invoke when user asks to audit tests, check test quality, find mock violations, review test effectiveness, or inspect test suites for over-mocking. Triggers automatic rewrites when quality gates fail.
tdd
Test-driven development workflow with philosophy guide - plan → write tests → implement → validate.