Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/xzawed/claude-grok-build-plugin/testsgit clone --depth 1 https://github.com/xzawed/claude-grok-build-pluginWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/xzawed/claude-grok-build-plugin/tests)<a href="https://agentmods.dev/commands/xzawed/claude-grok-build-plugin/tests"><img src="https://agentmods.dev/badge/commands/xzawed/claude-grok-build-plugin/tests.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00011 | $0.00212 |
| Opus 5 | $0.00005 | $0.00106 |
| Sonnet 5 | $0.00002 | $0.00042 |
| Haiku 4.5 | $0.00001 | $0.00021 |
Grade A, and why
tests scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Preset: use Grok for test backfill / expansion — a strong Grok fit (low risk, repetitive).
- Call
grok_auth_check. Ifok: false, showmessageand stop (guide/grok:setup). - Build an English
promptthat includes:- What to test (paths, functions, or "cover untested code in …")
- Constraints (framework, no flaky time/network, match existing style)
- Done criteria (tests run or at least compile; no production behavior change unless asked)
- Prefer
grok_build_verify(self-check) with absolutecwd. If the user wants a plan only, usegrok_build_planfirst. - For large or risky suites touching many packages, set
worktree: true. - Show
summary,filesChanged, andbilling. Review diffs; do not commit.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 15 lines · 11 tokens per session scan A 95000e4ce88c
tests is a command published in the GitHub repository xzawed/claude-grok-build-plugin (1 stars, last pushed yesterday), licensed MIT. It adds 11 tokens to every session and 212 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
psql-query
Run ad-hoc PostgreSQL analytics queries against dev/test database.
techdebt
Find and report technical debt in the codebase.
merge
Finish a PR properly: every check green, every review addressed — human and bot — then merge and clean up.
pr
Prepare and open a pull request the senior way: gate, template, scrubbed, everything visible.
spec
Spec-first design: a gap-closing interview that produces a complete spec, with a quality controller that blocks until every section is answered and every question resolved.
plan
Turn an approved spec into an implementation plan an engineer with zero context could execute — with a quality controller that blocks placeholders and hollow tasks.