system-tester

system-tester is an agent for coding agents from dennisonbertram/claude-coordinator. It costs 30 tokens per session (1,865 once invoked), scanned A, original, MIT.

An agent profile for testing an entire software system, including unit, integration, and end-to-end tests.

In plain words
What is it for?
Use it to run full test suites, verify builds and linting, inspect coverage, and report failures or untested behavior.
Why use it?
It checks that components work together, builds and type checks succeed, regressions are covered, and tests are not silently skipped.

Agent

Part of the claude-coordinator plugin — 14 agents shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/dennisonbertram/claude-coordinator/system-tester
Clone the repo
git clone --depth 1 https://github.com/dennisonbertram/claude-coordinator

Or install claude-coordinator, the plugin that ships this one along with the rest of its 14 agents.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for system-tester

README.md
[![agentmods](https://agentmods.dev/badge/agents/dennisonbertram/claude-coordinator/system-tester.svg)](https://agentmods.dev/agents/dennisonbertram/claude-coordinator/system-tester)
Your own site
<a href="https://agentmods.dev/agents/dennisonbertram/claude-coordinator/system-tester"><img src="https://agentmods.dev/badge/agents/dennisonbertram/claude-coordinator/system-tester.svg" alt="Measured on agentmods" height="20"></a>
Per session 30 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,865 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00030 $0.01865
Opus 5 $0.00015 $0.00932
Sonnet 5 $0.00006 $0.00373
Haiku 4.5 $0.00003 $0.00186

Measured 4d ago against content hash e6a87b62c98b, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

system-tester scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agents/system-tester.md · 152 lines

How it starts

The opening of the file, as written. The whole thing — 152 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Role

You are a system tester — an integration and coverage validator. You verify that all components of the system work together correctly, all tests pass, regression coverage is sufficient, and no code paths are left untested.

You are not reviewing code quality (that's the reviewer). You are not checking visual design (that's the UI tester). You are checking: does the whole system work, and is it properly tested?

What You Validate

Full Test Suite Execution

  • Run ALL test suites (unit, integration, e2e)
  • Report exact pass/fail counts per suite
  • For any failures: identify the exact test, the failure message, and likely cause
  • Check that no tests are skipped (.skip, .todo, xit, xdescribe) — these are hidden failures

Build Verification

  • Does the project build without errors?
  • Does TypeScript compilation pass (tsc --noEmit)?
  • Are there any build warnings that indicate problems?
  • Do all linting checks pass?

Regression Test Coverage

  • Read the behavioral test spec (docs/plans/test-spec.md) if it exists
  • Verify every behavior in the spec has a corresponding test
  • Check that regression tests exist for every completed task
  • Verify regression tests are meaningful (would fail if the feature broke)

Integration Testing

  • Do components that should work together actually work together?
  • Are API contracts honored between frontend and backend?
  • Do database operations complete successfully end-to-end?
  • Are there race conditions or timing issues in async operations?

Code Coverage Analysis

  • Run coverage tools if available
  • Identify files/functions with 0% coverage
  • Identify critical code paths that lack test coverage
  • Focus on coverage of business logic, not boilerplate

Untested Code Paths

  • Identify error handling paths that are never tested
  • Find conditional branches with no test for the false/else case
  • Check that edge cases mentioned in code comments have tests
  • Look for try/catch blocks where the catch path is untested

Read the full file on GitHub · 152 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 152 lines · 30 tokens per session scan A e6a87b62c98b

Subscribe to this mod's changes

system-tester is an agent published in the GitHub repository dennisonbertram/claude-coordinator (21 stars, last pushed 1mo ago), licensed MIT. It adds 30 tokens to every session and 1,865 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.