Testing

18,225 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

recalibrate-dash-p

554

ybouane/dash-p

Skill Claude CodeCodex

Recalibrate dash-p's recognition profile when a new Claude Code version ships. Drives the new TUI through a diverse SCENARIO BATTERY (short/long input, long output, code, markdown, tables, unicode, tool use), cross-checks every result against the session JSONL (ground truth), then fixes the profile/recognizer and…

not rated 19 2mo ago A 128 tokens

stably-sdk-setup

555

stablyai/agent-skills

Skill Claude CodeCodex

Expert setup assistant for the Stably Playwright SDK. Use this skill when installing Stably SDK in a new project, migrating from @playwright/test, or configuring Stably reporter for CI/CD. Triggers on tasks like "setup stably", "install stably sdk", or "configure playwright with stably".

not rated 19 4mo ago A 73 tokens

mcp-tester-guide

556

gleanwork/mcp-server-tester

Skill Claude CodeCodex

Reference guide for @gleanwork/mcp-server-tester — the Playwright-based testing and evaluation framework for MCP servers. Covers import paths, all 11 matchers, transport config, eval datasets, reporter setup, CLI commands, auth patterns, and common anti-patterns. Use when working with MCP server tests or evals.

not rated 19 yesterday A SkillSpector: warn 73 tokens original MIT

e2e-runner

557

JSK9999/ai-nexus

Agent Claude Code needs its repo

End-to-end testing specialist using Playwright. Use PROACTIVELY for generating, maintaining, and running E2E tests. Manages test journeys, quarantines flaky tests, uploads artifacts (screenshots, videos, traces), and ensures critical user flows work.

not rated 19 6mo ago A 59 tokens original Apache-2.0

cad-verification

558

rishigundakaram/cadquery-mcp-server

MCP server Claude CodeCodexCursor +2

MCP server "cad-verification" as configured in rishigundakaram/cadquery-mcp-server. Launched with /Users/rishigundakaram/.pyenv/shims/uv --directory /Users/rishigundakaram/Deskto.

not rated 19 +1 1y ago A tokens not measured

testing

559

manakuro/asana-clone-app

Skill Claude Code

Use this skill when writing, reviewing, or deciding where to place tests in the Next.js frontend package. Covers which project (unit / integration / storybook) to use for a given scenario, what NOT to duplicate across projects, and file placement conventions for this package.

not rated 19 4d ago A 56 tokens

plan-n-build

560

sam-agents/sam

Command Claude Code

SAM end-to-end composer - runs plan, then tdd for every story, then comprehensive docs. The one-shot PRD-to-working-code experience.

not rated 18 1mo ago A 34 tokens original MIT

belt

561

jfrog/agent-belt

Skill Claude CodeCodex needs its repo

Operate the belt CLI to evaluate headless coding agents (Claude Code, Cursor, Codex, Gemini, and others) end to end. Use when the user asks to write or run eval scenarios, compare agents, score outputs with rules or LLM judges, register a new agent adapter, interpret reports or benchmark cards, or set up evals in CI.…

not rated 18 5d ago A SkillSpector: pass 116 tokens original Apache-2.0

rashomon

562

shinpr/rashomon

Plugin Claude Code

Bundles 5 skills, 7 agents · 472 tokens together

Evaluate prompts and skills through parallel execution comparison in git worktrees. Prompt optimization with BP-001009 patterns, skill creation/update with quality grading, and blind A/B evaluation.

not rated 18 changed 3d ago A tokens not measured original MIT

verify

563

corveil/crow

Skill Claude Code

Drive CrowTelemetry's OTLP ingest end-to-end — boot the real receiver, POST OTLP JSON with curl, inspect the SQLite db.

not rated 18 today A SkillSpector: warn 30 tokens original Apache-2.0

JinNing6/Noosphere

Skill Claude CodeCodex

Verify Python packages, CLIs, MCP servers, and Agent plugins through the exact installed artifact and real runtime contract. Use when source tests pass but a published release may omit modules, expose broken entry points, depend on local paths, fail without optional credentials, or behave differently after public…

not rated 18 today A SkillSpector: warn 67 tokens original Apache-2.0

riwsky/iosef

Skill Claude CodeCodex needs its repo

Interaction with the iOS simulator using iosef, a CLI optimized for agent usage. Use when building or testing changes on the iOS Simulator — viewing the screen, tapping buttons, reading accessibility trees, finding elements by selector, asserting UI state, scripting multi-step test flows, installing and launching…

not rated 18 13d ago A SkillSpector: warn 100 tokens original MIT

pulser

566

TheStack-ai/pulser

Skill Claude Code

Diagnose and test Claude Code skills against Anthropic's 7 principles. Scans SKILL.md files, checks 8 rules (gotchas, description, allowed-tools, file-size, structure, frontmatter, conflicts, usage-hooks), classifies skill types, generates prescriptions, and runs eval tests. Use when checking skill quality, auditing…

not rated 18 2mo ago A 126 tokens original MIT

lastest

567

las-team/lastest

Skill Claude CodeCodex

Run visual regression tests, review screenshot diffs, and manage baselines on a Lastest instance via the @lastest/mcp-server MCP tools.

not rated 18 today A 33 tokens

ways-tests

568

aaronsb/claude-code-config

Skill Claude Code

Score way matching (embedding and BM25), analyze vocabulary, and validate frontmatter. Use when testing how well a way matches prompts, checking cosine similarity or BM25 scores, inspecting the embedding engine status, or validating way files.

not rated 18 5mo ago A 49 tokens original MIT

argos-cli

569

argos-ci/argos-javascript

Skill Claude CodeCodex

Operate Argos visual testing from the terminal with the argos CLI — inspect builds and snapshot diffs, submit reviews, request reviewers, post comments, inspect a test's flakiness and its recurring changes, ignore flaky test changes, configure a project and its contributors, manage a team's members, invites and email…

not rated 18 today A Socket: passSnyk: failSkillSpector: warn 143 tokens original MIT

hills

570

autolab-ai/hills

Skill Claude CodeCodex

Optimize anything through autonomous experimentation, against an evaluator you cannot grade yourself with. Use whenever the user wants to improve a number by iterating - raise accuracy, cut val loss, make training or inference faster, reduce latency or cost, beat a baseline, tune a pipeline or prompt, squeeze a…

not rated 19 10d ago A SkillSpector: warn 130 tokens original MIT

epic

571

epicsagas/epic-harness

Plugin Claude Code

Bundles 27 skills, 6 hooks · 997 tokens together

3 commands + 25 auto-trigger skills + self-evolving agent harness. Less to remember, more automation.

not rated 18 +1 changed 3d ago A tokens not measured original Apache-2.0

ever-works/directory-web-template

Skill Claude CodeCodex

Use when writing Playwright tests, fixing flaky tests, debugging failures, implementing Page Object Model, configuring CI/CD, optimizing performance, mocking APIs, handling authentication or OAuth, testing accessibility (axe-core), file uploads/downloads, date/time mocking, WebSockets, geolocation, permissions…

not rated 18 yesterday A 214 tokens AGPL-3.0

agent-evaluation

573

ihatesea69/kiro-kit

Skill Claude CodeCodex

Build evaluation harnesses for agents and chatbots — golden sets, deterministic tool-selection checks, LLM-as-a-Judge, Bedrock RAG evaluation jobs, CI gates, and online drift monitoring. Use when an agent needs a quality gate before merge or a quality alarm in production.

not rated 18 19d ago A SkillSpector: pass 61 tokens original MIT

goal-test

574

restarter/lets-workflow

Skill Claude Code

A local experiment for testing a goal command that keeps an AI coding session working until a stated condition is judged complete. It uses a separate language model to evaluate the conversation after each assistant turn.

not rated 17 3d ago A SkillSpector: warn 137 tokens original MIT

basic

576

permitio/permit-fastmcp

Cursor rule Cursor

When a soild fix - apply it without asking.

not rated 17 7mo ago A 0 tokens original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: