The best testing mods

104,027 mods in the catalogue carry this category. These 15 are the ones that pass every gate: original work, not a copy · security scan A or B · a licence that permits reading the source · an actively maintained repository people actually use. Order is the composite score — reputation, freshness, safety, originality, content and traction — not anyone's opinion.

obra/superpowers

Skill Claude CodeCodex

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

280k +583 today A 21 tokens original MIT

vercel/next.js

Skill Claude CodeCodex

Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…

142k yesterday A 170 tokens original MIT

auto-perf-optimize

03

microsoft/vscode

Skill Claude CodeCodex

Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.

190k yesterday A 62 tokens original MIT

playwright-cli

05

microsoft/playwright

Skill Claude CodeCodex

Automate browser interactions, test web pages and work with Playwright tests.

95k 3d ago A 19 tokens original Apache-2.0

microsoft/playwright

Skill Claude CodeCodex

Query Playwright CI test results from the aggregated DuckDB database. Answers questions about flaky tests, failure rates, slow tests, and per-run/SHA/PR results without hunting through GitHub artifacts.

95k 3d ago A 45 tokens original Apache-2.0

microsoft/playwright

Instructions file

Instructions for microsoft/playwright, covering monorepo packages, browser packages, tooling packages, key directories and build.

95k 3d ago A 1,902 tokens original Apache-2.0

microsoft/playwright

Agent Claude Code

Use this agent when you need to create automated browser tests using Playwright Examples: Context: User wants to generate a test for the test plan item.

95k 3d ago A 151 tokens original Apache-2.0

openai/openai-cookbook

Skill Claude CodeCodex

Bootstrap a new realtime eval folder inside this cookbook repo by choosing the right harness from examples/evals/realtimeevals, scaffolding prompt/tools/data files, generating a useful README, and validating it with smoke, full eval, and test runs. Use when a user wants to start a new crawl, walk, or run realtime eval…

76k 3d ago A 76 tokens original MIT

design-review

10

garrytan/gstack

Skill Claude CodeCodex

Designer's eye QA: finds visual inconsistency, spacing issues, hierarchy problems, AI slop patterns, and slow interactions — then fixes them. (gstack).

130k 2d ago A 36 tokens original MIT

nexu-io/open-design

Instructions file CodexOpenCode

Instructions for nexu-io/open-design, covering directory guide, core documentation index, workspace directories, inactive or placeholder directories and development workflow.

93k 2d ago A 9,822 tokens original Apache-2.0

material-ui-review

12

mui/material-ui

Skill Claude CodeCodex

Review the current diff for regressions, correctness bugs, tests, simplifications, and docs issues, scaling depth to a low/medium/high/xhigh/max effort level. Use ONLY when explicitly requested by name: the user runs /material-ui-review, writes $material-ui-review, or asks for "the Material UI review skill". Do NOT…

99k yesterday A 137 tokens original MIT

electron/electron

Skill Claude CodeCodex

Guide for performing Node.js version upgrades in the Electron project. Use when working on the roller/node/main branch to fix patch conflicts during e sync --3. Covers the patch application workflow, conflict resolution, analyzing upstream Node.js changes, building, running the Node.js test suite, and proper commit…

123k yesterday A 69 tokens original MIT

skill-creator

14

bytedance/deer-flow

Skill Claude CodeCodex

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

81k 2d ago A 64 tokens original MIT

dogfood

15

vercel-labs/agent-browser

Skill Claude CodeCodex

Systematically explore and test a web application to find bugs, UX issues, and other problems. Use when asked to "dogfood", "QA", "exploratory test", "find issues", "bug hunt", "test this app/site/platform", or review the quality of a web application. Produces a structured report with full reproduction evidence -…

42k 2d ago A 100 tokens original Apache-2.0

Want everything, not just what passes the gates? Browse the full category from any type hub, or search it. This shelf recomputes nightly from the measurements.