Testing

18,217 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

livekit-simulations

337

livekit/agent-skills

Skill Claude CodeCodex

Generate targeted test scenarios for a LiveKit voice or chat agent and run them as simulations — locally, from the agent's own code plus what the user wants stress-tested. Use whenever the user wants to "test my agent", "what should I test", "create/generate simulation scenarios", "make a sim test suite", "use lk…

not rated 67 +1 2mo ago A 183 tokens original MIT

skill-evaluator

338

HeshamFS/materials-simulation-skills

Skill Claude CodeCodex

Rigorously evaluate an Agent Skill end-to-end across ANY coding-agent CLI — verify its scripts emit the documented numbers (deterministic checks), test whether its description triggers on the right prompts, and measure whether an agent following the SKILL.md beats a no-skill baseline (with/without pass-rate delta…

not rated 66 +1 2mo ago A 181 tokens original Apache-2.0

adk-student-evaluator

339

mauripsale/doc-adk-training

Skill Claude CodeCodex

Acts as a student developer to evaluate ADK training modules. Follows a rigorous 5-step workflow to test labs and generate formal pedagogical reports. Use when asked to "evaluate", "test as a student", or "review" a training module.

not rated 66 +2 2d ago A 59 tokens

flutter-tester

340

Harishwarrior/flutter-claude-skills

Skill Claude CodeCodex

Use when creating, writing, fixing, or reviewing tests in a Flutter project. Covers unit tests, widget tests, integration tests, Riverpod provider testing, and Mockito mocking. Provides Given-When-Then patterns, layer isolation strategies, and test setup for GetIt, SharedPreferences, and FakeDatabase.

not rated 66 1mo ago A 65 tokens original MIT

mcp

341

mailtrap/mailtrap-mcp

MCP server Claude CodeCodexCursor +2

Official MCP Server for Mailtrap. Runs locally from the mcp-mailtrap npm package. Needs 2 environment variables to run.

not rated 65 3d ago A tokens not measured original MIT

test-writer

343

anchildress1/awesome-github-copilot

Skill Claude CodeCodex

Write, extend, or review tests in any codebase. Use this skill whenever the user asks to write tests, add test coverage, test a new feature, fix failing tests, or audit existing test files — regardless of language, framework, or project. Also trigger for "add tests for", "write tests for", "cover this with tests"…

not rated 65 3d ago A 130 tokens original MIT

skill-creator

344

Zhang-Henry/CoEvoSkills

Skill Claude CodeCodex

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

not rated 65 +1 16d ago A 64 tokens original Apache-2.0

BetterThanTomorrow/calva-backseat-driver

Skill Claude CodeCodex

Testing strategies for Calva Backseat Driver MCP tools. Use when: Testing Backseat Driver, validating tool updates, testing structural editing workflows, verifying REPL evaluation with who-tracking, testing shadow-cljs runtime discovery and targeting, testing output log filtering, testing load-file tool, smoke testing…

not rated 64 8d ago A 106 tokens original MIT

jmeter-mcp-server

346

juliodelimas/jmeter-mcp-server

MCP server Claude CodeCodexCursor +2

Stdio MCP server to build, run and read reports for JMeter test plans. Runs locally from the jmeter-mcp-server npm package.

not rated 64 +64 2d ago A tokens not measured original MIT

Kaggle/kaggle-skills

Skill Claude CodeCodex

Write, push, run, publish, and manage Kaggle Benchmark tasks using the kaggle CLI and the kaggle-benchmarks Python SDK. Use when the user wants to create or push a benchmark task (optionally with attached Kaggle datasets), run benchmarks against LLM models, check task/run status, stream or fetch execution logs…

not rated 63 +2 22d ago A 95 tokens original Apache-2.0

agent-inbox

348

gsd-build/agent-inbox

Skill Claude CodeCodex

Create temporary email inboxes and receive emails for testing auth flows, email verification, account confirmation, and any scenario where an AI agent needs to receive an email. Uses the agent-inbox MCP server with mail.tm + 1secmail fallback.

not rated 61 4mo ago A 54 tokens original MIT

api-tester

349

zeroclaw-labs/zeroclaw-skills

Skill Claude CodeCodex

Automated API testing with schema validation and load testing. Test REST and GraphQL APIs by sending requests, validating responses against schemas, checking edge cases, and reporting failures clearly. Use when the user wants to test APIs, validate endpoints, run load tests, or check schema compliance.

not rated 61 3d ago A 60 tokens

GitHub-Zero123/MCDevTool

Skill Claude CodeCodex

A workflow for testing Minecraft Bedrock Addons, Python 2 mods, interfaces, gameplay logic, and visual resources with MCDK-MCP. MCP is a way for an AI coding agent to call external testing tools.

not rated 61 5d ago A 0 tokens original BSD-3-Clause

node-testing

351

gleanwork/mcp-server

Cursor rule Cursor

Always prefer inline snapshot matching whenever possible. If necessary, have local normalization/sanitization functions to replace things that are noisy or unstable, like timestamps or error stacks.

not rated 61 2mo ago A 0 tokens original MIT archived

naveedharri/benai-skills

Agent Claude Code

Eval Agent for AutoResearch. Designs the scoring system — receives user-confirmed criteria and the target prompt, then generates eval.py + testcases.json (deterministic mode) or rubric.md + testcases.json (AI judge mode). The main agent never sees the eval artifacts in detail.

not rated 60 2d ago A 64 tokens original MIT

Zaoqu-Liu/ScienceClaw

Skill Claude CodeCodex

Create high-quality ToolUniverse skills following test-driven, implementation-agnostic methodology. Integrates tools from ToolUniverse's 1,264+ tool library, creates missing tools when needed using devtu-create-tool, tests thoroughly, and produces skills with Python SDK + MCP support. Use when asked to create new…

not rated 60 +1 5mo ago A 90 tokens original MIT

builder

354

kylehughes/apple-platform-build-tools-claude-code-plugin

Agent Claude Code

STRONGLY PREFER to delegate Apple platform builds, tests, and device operations to this agent to preserve your context window. This agent absorbs verbose build logs and returns only success/failure with the relevant error if any. Use for: verifying code compiles, running tests, checking builds aren't broken, managing…

not rated 59 6mo ago A 90 tokens original MIT

zhangchenglo/jmeter-function-developer-skills

Skill Claude CodeCodex

Develop Apache JMeter custom functions in Java with JavaFaker support. Use when creating random data generators (names, phones, emails, addresses), date/time functions, or custom logic for JMeter test plans.

not rated 59 7mo ago A 48 tokens

frontend-e2e

359

intellegix/intellegix-code-agent-toolkit

Command Claude Code

Perform a thorough live E2E test of a frontend web application using browser-bridge MCP tools. Tests every page, interactive element, form, navigation flow, responsive breakpoint, and accessibility rule through real browser interaction.

not rated 57 6d ago A 0 tokens original MIT

AmadeusITGroup/otter

Skill Claude CodeCodex

Write keyboard-only test variants and focus management checks using Playwright: navigation patterns, focus expectations, helpers, and file structure conventions.

not rated 57 yesterday A 34 tokens original BSD-3-Clause

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: