Testing mcp servers

28,585 tagged Testing, measured the same way as everything else here.

Browse within: accessibility 13Evaluation 12ai-testing 10qa 10Structured Output 9api-testing 9browser-automation 9token-efficiency 9AI Safety 7a11y 7ci 7code-quality 7regression-testing 7agent-evaluation 6

test-macos-app

217

openai/plugins

Command ✓ vendor

Run the smallest meaningful macOS test scope first and explain failures by category.

5.3k +48 5d ago A 0 tokens

create-skill-test

218

dotnet/skills

Skill Claude CodeCodex

Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml schema, fixture organization, and…

5.3k +17 today A 98 tokens original MIT

tdd

219

Galaxy-Dawn/claude-scholar

Command

Enforce test-driven development workflow. Scaffold interfaces, generate tests FIRST, then implement minimal code to pass. Ensure 80%+ coverage.

5.3k +37 6d ago A 28 tokens original MIT

fraimz

220

Devin-AXIS/iPolloWork

Command

Make fraimz for a flow — run the eval loop and output frame-by-frame proof (fraimz.html).

5.2k +122 yesterday A 24 tokens

test-writer

221

shareAI-lab/Kode-CLI

Agent

Specialized in writing comprehensive test suites. Use for creating unit tests, integration tests, and test documentation.

5.2k +2 6d ago A 25 tokens original Apache-2.0

spec-driven-eval

222

tech-leads-club/agent-skills

Skill Claude CodeCodex

Scores how completely an implementation fulfills a PRD/spec, case by case, and produces a single comparable final grade. Invoke only when explicitly named (e.g. run spec-driven-eval); do not auto-trigger. Use when benchmarking spec-driven implementations, grading acceptance criteria, evaluating whether a feature was…

5.1k +11 yesterday A 123 tokens

crabbox

223

openclaw/Peekaboo

Skill Claude CodeCodex

Use the Crabbox wrapper for OpenClaw remote validation across Linux, macOS, Windows, and WSL2, including delegated Blacksmith Testbox proof. Report the actual provider and id.

5.1k +22 2d ago B 43 tokens original MIT

dev

224

entireio/cli

Agent Claude Code

TDD Developer agent - implements features using test-driven development and clean code principles.

5.0k +7 today A 17 tokens original MIT

dev

225

entireio/cli

Command Claude Code

TDD Developer agent - implements features using test-driven development and clean code principles.

5.0k +7 today A 15 tokens original MIT

e2e

226

entireio/cli

Plugin Claude Code

E2E test triage, debugging, and fix implementation toolkit.

5.0k +7 today A tokens not measured original MIT

agent-integration

227

entireio/cli

Skill Claude CodeCodex

Run all three agent integration phases sequentially: research, write-tests, and implement using E2E-first TDD (unit tests written last). For individual phases, use /agent-integration:research, /agent-integration:write-tests, or /agent-integration:implement. Use when the user says "integrate agent", "add agent…

5.0k +7 today A 89 tokens original MIT

test-repo

228

entireio/cli

Skill Claude CodeCodex

Use this skill to test strategy changes against a fresh test repository. Invoke when the user asks to "test against a test repo", "validate the changes", or wants to verify session hooks, commits, and checkpoint creation work correctly.

5.0k +7 changed today A 50 tokens original MIT

kiln-prerelease-check

229

Kiln-AI/Kiln

Skill Claude CodeCodex

Run the Kiln pre-release smoke test suite plus the standard CI checks (checks.sh), diagnose every prerelease test that broke (and why), and write a clean readable report with recommended actions. Read-only — it never edits code. Use when the user wants to validate a release candidate, run prerelease tests, or asks for…

5.0k +2 today A 84 tokens

agent-device-evidence

230

Expensify/App

Skill Claude CodeCodex

Records iOS/Android native MP4 evidence for test/repro flows extracted from an Expensify GitHub PR or issue. Use when the user asks to "record the flow for PR.

5.0k +2 today A 43 tokens original MIT

agent-device

231

Expensify/App

Skill Claude CodeCodex

Drive iOS and Android devices for the Expensify App - testing, debugging, performance profiling, bug reproduction, and feature verification. Use when the developer needs to interact with the mobile app on a device.

5.0k +2 today A 45 tokens original MIT

peon-ping CLAUDE.md

233

PeonPing/peon-ping

Instructions file

Instructions for PeonPing/peon-ping, covering claude.md, commands, run all tests (requires bats-core: brew install bats-core), run a single test file and run a specific test by name.

5.0k +4 3d ago D 3,183 tokens original MIT

mockserver

234

mock-server/mockserver-monorepo

MCP server Claude CodeCodexCursor +2

MCP server "mockserver" as configured in mock-server/mockserver-monorepo. Runs in Docker (docker.io/mockserver/mockserver:7.6.0).

5.0k +4 today A tokens not measured original Apache-2.0

zep-eval-harness

235

getzep/zep

Skill Claude CodeCodex

Run and manage the Zep eval harness pipeline — document chunking, user ingestion, document ingestion, evaluation, graph inspection, and results analysis. Use when the user asks to run eval harness scripts, use the Zep eval harness, get terminal commands for eval harness operations, chunk documents, ingest users or…

4.9k +4 today A 139 tokens original Apache-2.0

harbor AGENTS.md

236

harbor-framework/harbor

Instructions file CodexOpenCode

AGENTS.md instructions for harbor-framework/harbor, covering claude.md - harbor framework, contributing, project overview, quick start commands and install.

4.9k +72 yesterday A 4,064 tokens original Apache-2.0

create-adapter

237

harbor-framework/harbor

Skill Claude CodeCodex

Scaffold a new Harbor benchmark adapter by running harbor adapter init and then guide implementation using the Adapters Agent Guide as the authoritative spec.

4.9k +72 yesterday A 34 tokens original Apache-2.0

rewardkit

238

harbor-framework/harbor

Skill Claude CodeCodex

Write Harbor task verifiers using Reward Kit. Use when creating or editing a task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing verifiers that produce a reward score.

4.9k +72 changed today A 46 tokens original Apache-2.0

99 AGENTS.md

239

ThePrimeagen/99

Instructions file CodexOpenCode

AGENTS.md instructions for ThePrimeagen/99, covering testing and e2e / integration style testing.

4.8k +1 2mo ago A 374 tokens

agent-release-gate

240

Agenta-AI/agenta

Skill Claude CodeCodex

Run the agent release gate — a portable, wire-level QA harness for the agent runtime. Drives the same product endpoint the playground drives and asserts on the SSE frame stream and real side effects, never on model prose, so it works against any deployment (cloud or self-hosted) from three env vars. Use before an…

4.7k +28 changed today A 126 tokens

At most 3 mods per repository are shown here — the rest are on their repository pages: