testing-specialist

testing-specialist is an agent for Claude Code from khill1269/servalsheets. It costs 59 tokens per session (3,321 once invoked), scanned A, original, MIT.

A testing assistant for ServalSheets that plans and creates different kinds of automated tests. It uses TDD, where tests are written before implementation, and BDD, where behaviour is described from the user's perspective.

In plain words
What is it for?
Add tests for new features, reproduce and prevent bugs, check critical code paths, and measure whether changes survive deliberately introduced faults.
Why use it?
It helps detect bugs and missing coverage across unit, integration, property-based, failure-injection, and live API tests.

Agent for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/khill1269/servalsheets/testing-specialist
Clone the repo
git clone --depth 1 https://github.com/khill1269/servalsheets

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for testing-specialist

README.md
[![agentmods](https://agentmods.dev/badge/agents/khill1269/servalsheets/testing-specialist.svg)](https://agentmods.dev/agents/khill1269/servalsheets/testing-specialist)
Your own site
<a href="https://agentmods.dev/agents/khill1269/servalsheets/testing-specialist"><img src="https://agentmods.dev/badge/agents/khill1269/servalsheets/testing-specialist.svg" alt="Measured on agentmods" height="20"></a>
Per session 59 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 3,321 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00059 $0.03321
Opus 5 $0.00030 $0.01661
Sonnet 5 $0.00012 $0.00664
Haiku 4.5 $0.00006 $0.00332

Measured 3d ago against content hash 8e44e28ae518, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

testing-specialist scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/agents/testing-specialist.md · 476 lines

How it starts

The opening of the file, as written. The whole thing — 476 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are a Testing Specialist focused on comprehensive, efficient test coverage for ServalSheets.

Your Expertise

Testing Infrastructure:

  • Test runner: Vitest (threads mode, CI=4 workers, local=8 workers)
  • Coverage: v8 provider — 50% lines/functions, 30% branches
  • Mutation testing: Stryker (9 critical files, threshold 60%)
  • Property-based: fast-check (8 test files)
  • Chaos: Failure injection (4 test files)
  • Live API: Real Google Sheets integration (40 test files, 3 tiers)

ServalSheets Test Pyramid (613 total test files):

Tier Dir Files Command
Audit tests/audit/ 5 npm run audit:coverage/perf/memory
Contract tests/contracts/ 41 npm run test:fast (included)
Handler tests/handlers/ 73 npm run test:fast (included)
Services tests/services/ 81 npm run test:services
Integration tests/integration/ 20 npm run test:integration
Compliance tests/compliance/ 15 npm run test:compliance
Property tests/property/ 8 npm run test:run tests/property
Chaos tests/chaos/ 4 npm run test:run tests/chaos
Packages tests/packages/ 38 npm run test:mcp-http-task-contract
Snapshots tests/snapshots/ 1 npm run test:snapshots
Live API tests/live-api/ 40 npm run test:live:smoke/nightly
E2E tests/e2e/ 10 npm run test:run tests/e2e
Security tests/security/ 8 npm run test:run tests/security

Live API tiers:

npm run test:live:smoke         # Quick subset (~10 min, smoke config)
npm run test:live:nightly       # Full suite (all 40 files, nightly config)
npm run test:live:optimizations # Performance optimization tests
npm run test:live:full          # Alias for nightly

Core Responsibilities

1. Test Strategy Design

For every new feature, create:

## Test Strategy: [Feature Name]

### Test Pyramid

- **Unit Tests** (70%): Fast, isolated, no dependencies
- **Integration Tests** (20%): Component interactions
- **E2E Tests** (10%): Full workflows

### Coverage Goals

- **Critical paths:** 100% (MUST be tested)
- **Error handling:** 100% (all error codes)
- **Happy paths:** 100%
- **Edge cases:** 95%

### Test Types Needed

1. ✅ Unit tests for pure functions
2. ✅ Handler tests for business logic
3. ✅ Contract tests for schemas
4. ✅ Property-based tests for invariants
5. ✅ Chaos tests for failure modes
6. ✅ Benchmark tests for performance

Read the full file on GitHub · 476 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 476 lines · 59 tokens per session scan A 8e44e28ae518

Subscribe to this mod's changes

testing-specialist is an agent published in the GitHub repository khill1269/servalsheets (0 stars, last pushed today), licensed MIT. It adds 59 tokens to every session and 3,321 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other agents, from other repositories

context-manager

Use this agent when you need to manage context across multiple agents and long-running tasks, especially for projects exceeding 10k tokens. This agent is essential for coordinating complex multi-agent workflows, preserving context across sessions, and ensuring coherent state management throughout extended development…

czlonkowski/n8n-mcp · 0 tokens

chainaware-token-launch-auditor

Audits a new token launch for launchpads by combining rug pull detection on the contract with fraud and behavioral analysis on the deployer wallet. Returns a composite Launch Safety Score, a APPROVED / CONDITIONAL / REJECTED listing verdict, a public-facing safety badge, and specific conditions the launchpad should…

ChainAware/behavioral-prediction-mcp · 247 tokens

Plan

Research and outline multi-step plans for zen analysis improvements.

Anselmoo/mcp-zen-of-languages · 13 tokens

issue-tracker

Issues and PRDs for this repo live as GitHub issues. Use the gh CLI for all operations.

MiaoDX/roboclaws · 0 tokens

review

Pre-PR code review against the project's gates and cross-cutting contracts — read-only, run before any external reviewer.

hybridindie/comfyui_mcp · 23 tokens

comment-fixer

Scans source files and fixes code comments; adds missing one-line JSDoc, improves existing JSDoc, and cleans up inline comments (WHY not WHAT, removes obvious or stale ones). Defaults to recently changed files; prompt with full for a whole-src sweep. Use when asked to clean up, fix, or standardize comments.

timohaa/scopewalker-mcp · 74 tokens