Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/davistroy/claude-marketplaceWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/davistroy/claude-marketplace/test-project.eval)<a href="https://agentmods.dev/commands/davistroy/claude-marketplace/test-project.eval"><img src="https://agentmods.dev/badge/commands/davistroy/claude-marketplace/test-project.eval/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/commands/davistroy/claude-marketplace/test-project.eval"><img src="https://agentmods.dev/badge/commands/davistroy/claude-marketplace/test-project.eval.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00000 | $0.00610 |
| Opus 5 | $0.00000 | $0.00305 |
| Sonnet 5 | $0.00000 | $0.00122 |
| Haiku 4.5 | $0.00000 | $0.00061 |
Grade A, and why
test-project.eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 101 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Eval: /test-project
Purpose
Ensures 90%+ test coverage, runs all tests with sub-agents, fixes failures, then creates a PR. This is a comprehensive test-and-ship workflow. Evals should run against a project that has tests, not the marketplace repo itself.
Fixtures
None — requires a code project with tests. Use a separate Node.js or Python project for eval.
Setup
# Create a minimal test project with some failing tests
mkdir /tmp/eval-test-project && cd /tmp/eval-test-project
git init && git checkout -b main
npm init -y
npm install --save-dev jest
# Create a simple module with tests (some passing, some failing)
Test Scenarios
S1: Project with passing tests
Setup: Project with a test suite where all tests pass.
Invocation: /test-project
Must:
- Discovers and runs existing tests
- Reports test results (pass count, fail count, coverage percentage)
- If coverage < 90%, adds tests to close the gap
- Creates a PR after tests pass
Should:
- Uses sub-agents to run tests in parallel when possible
- Provides a summary of what was tested
Must NOT:
- Delete existing tests to "fix" coverage
- Skip tests to get a passing suite
- Push to main branch
S2: Project with failing tests
Setup: Project with one deliberately failing test.
Invocation: /test-project
Must:
- Identifies the failing test(s)
- Attempts to fix the failing tests
- Re-runs tests after fix to verify resolution
- Does not create a PR until tests pass
Should:
- Reports what the test failure was and how it was fixed
Must NOT:
- Delete failing tests as a "fix"
- Create a PR with failing tests
S3: --auto-merge flag
Invocation: /test-project --auto-merge (in isolated test project)
Must:
- Creates PR and merges it after all tests pass
S4: Prerequisite — not on main
Setup: Be on main branch.
Invocation: /test-project
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 101 lines · 0 tokens per session scan A 4272f5abf0f6
test-project.eval is a command published in the GitHub repository davistroy/claude-marketplace (5 stars, last pushed yesterday), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 610 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
verify
Run repository verification using the verification-loop skill.
qa-changes
This skill should be used when the user asks to "QA a pull request", "test PR changes", "verify a PR works", "functionally test changes", or when an automated workflow triggers QA validation of code changes. Provides a structured methodology for setting up the environment, exercising changed behavior, and reporting…
test-coverage
Analyze test coverage and identify the highest-value gaps to fill.
tdd
A command that follows test-driven development (TDD), a method where you write tests before the code they check. It moves through writing a failing test, adding the smallest implementation, and then improving the code.
check-dev
Type-check a Z specification with fuzz.
laravel-playwright
E2E Playwright patterns; use the laravel:e2e-playwright skill exactly as written.