Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/The-Artificer-of-Ciphers-LLC/skills-from-the-artificerWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/the-artificer-of-ciphers-llc/skills-from-the-artificer/test)<a href="https://agentmods.dev/commands/the-artificer-of-ciphers-llc/skills-from-the-artificer/test"><img src="https://agentmods.dev/badge/commands/the-artificer-of-ciphers-llc/skills-from-the-artificer/test/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/commands/the-artificer-of-ciphers-llc/skills-from-the-artificer/test"><img src="https://agentmods.dev/badge/commands/the-artificer-of-ciphers-llc/skills-from-the-artificer/test.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00023 | $0.00512 |
| Opus 5 | $0.00012 | $0.00256 |
| Sonnet 5 | $0.00005 | $0.00102 |
| Haiku 4.5 | $0.00002 | $0.00051 |
Grade A, and why
test scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Run the project's tests on all configured platforms (Mac local + Linux Docker + Windows Docker) and analyze the diff.
Steps
-
Run
gsd-test-both. It launchesgsd-test-local,gsd-test, and (if configured)gsd-test-windowsin parallel, writes JSON Lines results to per-invocation files in/tmp/, then prints a comparison summary. Windows is skipped with a notice if~/.config/gsd-test/windows-hostsis empty. -
Read the JSON Lines files. Parse and categorize failures into:
- All fail (real bugs — same failure on every platform)
- Mac-only fail (passes on Docker platforms — Mac-specific issue)
- Linux-only fail (passes on Mac + Windows — Linux/container-specific issue)
- Windows-only fail (passes on Mac + Linux — Windows-specific issue)
- Missing on one platform (test discovery differs)
-
Summarize for me:
- Counts per category
- For each failure type, list up to 10 with file + test name + first line of the error message
-
If anything is platform-specific, call those out as the most interesting findings.
-
If
gsd-test-bothexited non-zero, surface the stderr tails it printed.
Useful jq one-liners
# All failures on Linux:
jq -c 'select((.type=="test:fail") or (.type=="test_event" and .kind=="fail")) | {file: (.data.file // .file), name: (.data.name // .name)}' /tmp/gsd-test-docker-*.jsonl
# Test names grouped by file (most-broken files first):
jq -r 'select((.type=="test:fail") or (.type=="test_event" and .kind=="fail")) | (.data.file // .file)' /tmp/gsd-test-docker-*.jsonl | sort | uniq -c | sort -rn
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 39 lines · 23 tokens per session scan A ba89aa6471c9
test is a command published in the GitHub repository The-Artificer-of-Ciphers-LLC/skills-from-the-artificer (4 stars, last pushed 10d ago), licensed MIT. It adds 23 tokens to every session and 512 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
red-team
Stress-test a plan, strategy, PRD, or launch with a room of hostile expert personas before you commit.
test-cases
Use when writing test cases, generating tests, supplementing test coverage, or improving test completeness — auto-scans project, designs multi-dimensional test cases with coverage matrix.
qa-guide
Generate a complete, prioritized test plan for a feature.
edge-case-finder
Find edge cases for a feature or requirement.
ship-an-mcp-server
Workflow recipe — make your product agent-usable by chaining 4 skills, spec to pricing.
rescue-an-account
Workflow recipe — diagnose an at-risk customer and build the full save play through to renewal by chaining 4 skills.