Testing

18,481 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.

llm-eval

1321

pnakhat/qa-ai-repo

Skill Claude CodeCodex

Author LLM/RAG/agent evaluation suites in DeepEval that prove a feature is correct with gated numbers, not vibes. Use when asked to "eval an LLM", "test a prompt", "measure RAG quality", "check for hallucination", "score answer relevancy", "verify tool calls", or gate a release on model output quality. Ships the…

not rated 2 2mo ago A 194 tokens original MIT

wtalaat78/Blue-Team-Skills

Cursor rule Cursor

Comprehensive AppSec & authorized security testing workflow (OWASP Top 10, STRIDE, CVSS v3.1, 11-domain scoring, Pre-Fix reports, and SOC detection rules).

not rated 2 yesterday A 414 tokens

process

1323

origammi-digital/agents-marketplace

Plugin Claude Code

Bundles 4 skills, 4 agents · 399 tokens together

Engineering process discipline: how the work is conducted, not who does it. Debugging, TDD, planning, and verification — the workflow skills that resist being skipped under pressure.

not rated 2 2d ago A tokens not measured

Stramp/Lost-Mines-Phandelver

Cursor rule Cursor

CRITICAL RULE - Test-Driven Development (TDD) methodology. Write tests BEFORE implementation for critical code.

not rated 2 9mo ago A 4,668 tokens original MIT

development-workflow

1325

kurisu-dotto-komu/next-agentic-coding

Cursor rule Cursor

Cursor rule "development-workflow" from kurisu-dotto-komu/next-agentic-coding, covering development workflow standards, development server, git operations, documentation creation and testing and linting commands.

not rated 2 9mo ago A 439 tokens

Xclaw-bot/benchmark-task-authoring

Skill Claude CodeCodex

Design, red-team, ship and debug hard Terminal-Bench 2 / Harbor benchmark tasks: the measured laws for what makes agents actually fail, the kill-list of dead task shapes, and how to clear all 17 review stages in one push instead of three. Use for benchmark task slots, TB2/Harbor tasks, task.toml, instruction.md, task…

not rated 2 28d ago A 207 tokens original MIT

uvvr

1329

divo12/UVVR

Cursor rule Cursor

Design executable evals for AI outcomes that are open-ended or only partly verifiable.

not rated 2 13d ago A 16 tokens original MIT

sage

1330

gustavobarbosab/sage

Skill Claude CodeCodex

SAGE — spec-first, AI-assisted development workflow using SDD, BDD, and Harness Engineering. Use this skill whenever the user types /sage, wants to write a spec, generate behavior scenarios, generate code from a spec, export a PR description or documentation, manage a project harness, or mentions spec.md, behavior.md…

not rated 2 2mo ago A 122 tokens original MIT

uvs-qa

1331

utsavanand/uv-suite

Skill Claude Code

Browser-based QA: exercises the running app via Playwright MCP, captures console errors and visual evidence, optionally fixes source bugs with atomic commits and generates regression tests. Three tiers (quick / standard / exhaustive). Writes uv-out/qa-state.md so /uvs-commit and /ship can detect completion and read…

not rated 2 3mo ago A 71 tokens original MIT

premouseking/Mentora

Skill Codex

A local development check for the Mentora application, which runs across API, web, and desktop parts. It starts the project on a Windows computer and verifies that the parts work together.

not rated 2 2mo ago A 70 tokens

STAGE_S6_TEST_CASES

1333

SilenceInsect/AIDocxWorkFlow

Cursor rule Cursor

触发方式:/aidocx-s6-test-cases 或粘贴 S5 testpoints.json.

not rated 2 1mo ago A 98 tokens original MIT

java-backend-test-ops

1334

LSRabbit6/cursor-genesis

Skill Claude CodeCodex

A set of practices for running Java backend integration tests with Maven, Spring Boot, Testcontainers, Docker Desktop, and MySQL containers. It covers container resource issues, Spring Boot null-safety rules, and changes to shared test base classes.

not rated 2 26d ago A 147 tokens

forge

1335

yash-dev007/THE-FORGE

Cursor rule Cursor

THE FORGE v4 — self-driving loop via forge-cycle.sh, per-repo Quintet, 3-run EVAL median, Obsidian sync.

not rated 2 4mo ago A 0 tokens original MIT

scala-weaver-test

1336

sanssushi/skills

Skill Claude CodeCodex

Use when writing or modifying Scala 3 tests with weaver-cats, weaver-discipline, or weaver-scalacheck; covers suites, effects, expectations, resources, and laws.

not rated 2 1mo ago A 43 tokens original MIT

electron-testing

1337

Sovea/skills

Skill Codex

Autonomously plan, run, and assess evidence-driven tests for Electron applications. Use after implementing or refactoring Electron behavior, when reproducing an issue, or when validation crosses main, preload, renderer, IPC, multi-window, lifecycle, packaged-runtime, or native desktop boundaries. Select the smallest…

not rated 2 1mo ago A 99 tokens

web-ui-smoke

1338

avelrl/skills

Skill Claude CodeCodex

Use when Codex needs to run a local web app in a real browser, click/fill/press, wait for visible UI changes, save screenshots, and inspect console/page/request errors. Best default for browser smoke tests across repos.

not rated 2 4mo ago A 52 tokens original MIT

artilleryio/agent-skills

Skill Claude Code

Set up Artillery load testing for any project. Detects package manager and project type, creates a TypeScript test script (HTTP or Playwright browser), configures Artillery Cloud, and provides the run command. Use when the user wants to add load testing, performance testing, or browser-based load testing to their…

not rated 2 6mo ago A 73 tokens

browser-automation

1340

raywongstudy/agent_skills

Skill Claude CodeCodex

A browser-automation skill that uses ChromeDriver and Selenium to control a real Chrome browser for testing web pages.

not rated 2 1mo ago A 113 tokens

KarhouTam/agent-skills

Skill Claude CodeCodex

Orchestrate PyTorch test file refactoring to decouple tests from specific hardware accelerators. Use this skill whenever the user asks to refactor, decouple, or reorganize a PyTorch test file for cross-accelerator compatibility, or when they ask to apply the test decoupling workflow to a specific test file. Triggers…

not rated 2 17d ago A 176 tokens

playwright-testing

1342

dtinth/agent-skills

Skill Claude CodeCodex

Playwright testing. Use this skill to write and run automated tests for web applications using Playwright.

not rated 2 6mo ago A 24 tokens

code-rigor-check

1343

alok-19/agent-skills

Skill Claude CodeCodex

Runs a structured rigor check over code before it ships, with a specific mode for AI-generated output. Covers problem framing, edge cases, failure modes, explainability, and AI-specific failure patterns — hallucinated APIs, plausible-but-wrong library behavior, tests that mirror the code, defensive scaffolding that…

not rated 2 19d ago A 221 tokens

tdd-workflow

1344

dennishenle/agent-skills

Skill Claude CodeCodex

Use this skill when writing new features, fixing bugs, or refactoring code. Enforces test-driven development with 80%+ coverage including unit, integration, and E2E tests.

not rated 2 5mo ago A 43 tokens

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: