Testing

27,851 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

skillgrade-graders

193

mgechev/skillgrade

Skill Claude CodeCodex

Authors deterministic and LLM rubric graders for skillgrade evaluations. Use when creating scoring scripts, writing evaluation rubrics, or combining multiple graders with weighted scoring. Don't use for setting up eval pipelines, configuring eval.yaml defaults, or general test writing.

not rated 700 +8 10d ago A 54 tokens original MIT

rewardkit

194

zli12321/LHTB

Skill Claude CodeCodex

Write Harbor task verifiers using Reward Kit. Use when creating or editing a task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing verifiers that produce a reward score.

not rated 697 +5 9d ago A 46 tokens copy · 91% Apache-2.0

naver/fixture-monkey

Instructions file Claude Code

A Korean-language guide for working on Fixture Monkey, a Java and Kotlin library for generating test data. It describes project rules, folder roles, build commands, and testing requirements.

not rated 697 12d ago A 1,379 tokens original Apache-2.0

shotgun CLAUDE.md

196

shotgun-sh/shotgun

Instructions file Claude Code

Instructions for shotgun-sh/shotgun, covering claude code instructions for shotgun, evals, writing eval test cases, architecture documentation and commit message convention.

not rated 682 +2 3mo ago A 3,724 tokens original MIT

shinpr/claude-code-workflows

Agent Claude Code

Generates integration/E2E test skeletons from Design Doc ACs using ROI-based selection and journey-based E2E reservation. Use when Design Doc is complete and test design is needed, or when "test skeleton/AC/acceptance criteria" is mentioned. Behavior-first approach for minimal tests with maximum coverage.

not rated 679 +4 2d ago A 68 tokens original MIT

rayfish CLAUDE.md

198

rayfish/rayfish

Instructions file Claude Code

Claude Code instructions for rayfish/rayfish, covering rayfish, build & test, never, conventions and git.

not rated 663 +9 today A 1,014 tokens original MPL-2.0

eval-validity-review

199

UKGovernmentBEIS/inspect_evals

Skill Claude Code

Review a single evaluation's validity — whether its claims hold up, whether its name is accurate, whether samples can be both succeeded and failed at, and whether scoring measures ground truth. Use when user asks to check validity of an eval, or as part of the Master Checklist workflow. Do NOT use for code quality or…

not rated 658 +8 2d ago A 83 tokens original MIT

obsidian-e2e

200

uphy/obsidian-reminder

Skill Claude Code

An automated end-to-end test workflow for a running Obsidian application, controlled through Chrome DevTools Protocol. Obsidian is a note-taking app, and end-to-end testing checks behavior across the real application rather than only individual functions.

not rated 657 +3 4d ago A 140 tokens original MIT

wshobson/maverick-mcp

Instructions file CodexOpenCode

AGENTS.md instructions for wshobson/maverick-mcp, covering repository guidelines, project overview, project structure, documentation map and build, test, and development commands.

not rated 656 +3 4d ago A 1,422 tokens original MIT

callstackincubator/rozenite

Agent Claude Code

Rozenite is a development tool. Whatever a project wires into its bundler config, a release build must ship none of our code. @rozenite/test-utils provides the bench that proves it, and every plugin owns a Vitest suite in src/tests/release-bundle.test.ts that uses it.

not rated 656 +3 yesterday A 0 tokens original MIT

ClawBench AGENTS.md

204

TIGER-AI-Lab/ClawBench

Instructions file CodexOpenCode

AGENTS.md instructions for TIGER-AI-Lab/ClawBench, covering clawbench -- agent context, what this is, project structure, setup and 2. configure at least one model.

not rated 649 +40 changed 2d ago A 1,789 tokens original Apache-2.0

faker CLAUDE.md

205

jaswdr/faker

Instructions file Claude Code

Claude Code instructions for jaswdr/faker, covering claude.md, repository overview, core architecture, common commands and testing.

not rated 644 +1 1mo ago A 511 tokens original MIT

principal-qa-engineer

206

SixHq/Overture

Agent Claude Code

Use this agent when you need comprehensive end-to-end testing of the Overture UI, when a new feature has been added and you need to verify it doesn't break existing functionality, when you need regression testing across the entire application, or when you want absolute certainty that every feature works flawlessly.…

not rated 637 +2 6mo ago A 448 tokens original MIT

antfu

207

oyjt/uniapp-vue3-template

Skill Claude CodeCodex

Anthony Fu's {Opinionated} preferences and best practices for web development.

not rated 626 3mo ago A 17 tokens original MIT

michaelshimeles/skills

Skill Claude CodeCodex

Records visual proof while testing UI behavior — the agent tests the app hands-on via computer use while a screen recording with structured test/assertion annotations captures the session — then posts the video and a results summary to the PR and tracker issue. Use whenever a change needs verifiable evidence that it…

not rated 608 +278 changed 2d ago A 93 tokens

nw-distill

209

nWave-ai/nWave

Skill Claude Code

Acceptance test creation methodology for the DISTILL wave. Domain knowledge for the acceptance designer agent: port-to-port principle, prior wave reading, wave-decision reconciliation, graceful degradation, and document back-propagation.

not rated 605 +3 6d ago A 46 tokens original MIT

berabuddies/Semia

Skill Claude CodeCodex

Download YouTube videos in various formats and qualities. Use when you need to save videos for offline viewing, extract audio, download playlists, or get specific video formats.

not rated 595 4d ago B 38 tokens original Apache-2.0

aislop AGENTS.md

211

scanaislop/aislop

Instructions file CodexOpenCode

AGENTS.md instructions for scanaislop/aislop, covering ai agent instructions for aislop, what is aislop?, build & test commands, cross-platform scripts and writing conventions.

not rated 595 +7 5d ago C 1,785 tokens original MIT

add-benchmark

212

allenai/vla-evaluation-harness

Skill Claude Code

Add a new simulation benchmark to the VLA evaluation harness. Use this skill whenever the user wants to integrate, create, or add a new benchmark or simulation environment — e.g. 'add ManiSkill3', 'integrate OmniGibson', 'hook up a new sim'. Also use when they ask how benchmarks are structured or want to understand…

not rated 588 +15 4d ago A 78 tokens original Apache-2.0

solid

213

ramziddin/solid-skills

Skill Claude CodeCodex

Use this skill when writing code, implementing features, refactoring, planning architecture, designing systems, reviewing code, or debugging. This skill transforms junior-level code into senior-engineer quality software through SOLID principles, TDD, clean code practices, and professional software design.

not rated 582 +4 changed 2d ago A 56 tokens

pr-reviewer

214

pyuvm/pyuvm

Skill Claude CodeCodex

Evaluates GitHub Pull Requests against a Test Sufficiency Matrix and Intent Realization Alignment, or provides a high-level summary of all open PRs in the repository.

not rated 569 +1 4d ago A 37 tokens

johnpapa/vscode-angular-snippets

Instructions file CodexOpenCode

AGENTS.md instructions for johnpapa/vscode-angular-snippets, covering angular snippets for vs code — agent guide, repository structure, tech stack, build & run and test in browser (vscode.dev mode).

not rated 567 2d ago A 1,584 tokens original MIT

test-writer

216

vinilana/dotcontext

Agent Claude Code

Write comprehensive unit and integration tests.

not rated 565 +5 1mo ago A 10 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: