Skill Claude CodeCodex
Run FixThis release and trust-loop smoke checks.
11,777 tagged Testing, measured the same way as everything else here.
Browse within: LLM 180agents 133agentic-ai 129cli 101ai-coding 99agent 80skills 72javascript 57agent-browser 54openai 51ai-testing 46agentic-workflow 41agent-orchestration 40claude-code-plugin 40
Skill Claude CodeCodex
Run FixThis release and trust-loop smoke checks.
Skill Claude CodeCodex needs its repo
Argus AI-powered QA harness — Chrome DevTools MCP reference for browser automation, accessibility, performance, security, and debugging.
Skill Claude CodeCodex
· Audit AI-generated code slop: hallucinated APIs, over-abstraction, duplicate code, test theater, noisy comments. Triggers: 'slop', 'AI-generated code', 'cleanup', 'overengineered'. Not for prose (use anti-ai-prose).
Skill Claude CodeCodex
The only skill you need to build and ship world-class software — from idea to production. Covers the COMPLETE lifecycle: architecture design, domain modeling, service boundaries, data architecture, implementation with TDD, distinctive frontend design (anti-AI-slop), exhaustive QA across all layers, security hardening…
Skill Claude CodeCodex needs its repo
Playwright E2E testing patterns, Page Object Model, configuration, CI/CD integration, artifact management, and flaky test strategies.
Skill Claude CodeCodex
Build a new feature using strict test-driven development — UNDERSTAND → DESIGN → RED → VERIFY RED → GREEN → VERIFY GREEN → REFACTOR → VALIDATE → REPEAT. Use when building a new feature with TDD, when the user says "TDD this", "tests first", "red green refactor", or asks for feature work that needs to be verifiable…
Skill Claude CodeCodex
Use when a coding agent must turn a software goal, bug, issue, or refactor into a small, verified, independently reviewed code change.
chuanzige/claude-code-main-skills
Skill Codex
Use when verifying implementation work and the failure mode is superficial approval, code-reading in place of execution, or over-trusting a passing test suite without trying to break the change.
flyingsquirrel0419/squirrel-skill
Skill Claude CodeCodex
Full-cycle software development agent: plans, builds, tests, lints, fixes bugs, and writes production-grade README docs. ALWAYS use this skill when the user wants to: build or scaffold a new project, add features to existing code, fix bugs, improve code quality (lint, format, refactor), write or improve a README, add…
Skill Claude CodeCodex needs its repo
Render-and-measure receipts for any HTML page your factory builds — the render half of the design quality gate. Engineer runs it to screenshot every screen size and MEASURE what a source read or a single screenshot only guesses at: horizontal overflow, computed type sizes, tap-target sizes, safe-area presence, mono…
burugo/behavior-driven-development
Skill Claude CodeCodex
Use when implementing any feature or bugfix, before writing implementation code. Triggered by keywords: TDD, BDD, test-first, test-driven, write-test-first.
Skill Codex
AI-powered development system that takes a project from raw idea to hardened production. Fuses artifact discipline, execution engine, quality enforcement, and team intelligence into one unified workflow. Triggers on: "god mode", "god init", "god prd", "god arch", "god roadmap", "god stack", "god repo", "god build"…
Skill Claude CodeCodex needs its repo
Write skillprobe YAML tests for LLM skills. Use when asked to write tests, create test scenarios, test a skill, generate skillprobe tests, or check whether a skill activates correctly.
Skill Claude CodeCodex
Run Archetype AI's managed Task Verification (TVA) agent over the Agents API — upload a recording AND a reference procedure (an SOP), create a bundle from the tva blueprint, run it, poll, download a per-step PASSED / FAILED / MISSING verdict per step. Use when the user has a recording of work that should have followed…
andreferraro/skill-suite-tests
Skill Codex
Analisa riscos e cria, adapta, executa e valida testes automatizados em projetos existentes. Use quando o usuário pedir testes para uma tela, fluxo, regra, serviço, API, banco, evento, integração, bug ou atributo de qualidade e esperar código integrado à stack, execução e evidências reais.
Skill Codex
Design, install, run, and report deterministic DeepSeek Harness tool-failure experiments. Use when a user wants to prove retry or fallback behavior, timeout or cooperative cancellation, policy-denial handling, blocked-result recovery, Code Mode nested-call resilience, or CI evidence for a DSH agent/plugin. Complete…
Skill Claude CodeCodex
Configure and debug FinFocus intelligent plugin routing. Use when setting up multi-plugin routing, configuring priority and fallback rules, writing pattern matching for resource types, testing route selection, or debugging why a plugin isn't receiving cost queries. Triggers on: "configure routing", "plugin priority"…
Skill Codex
A test-planning skill that turns a specification's requirements and acceptance conditions into observable results, test scenarios, evidence, and proof obligations. TDD means writing tests or defining expected behavior before implementation.
Skill Claude CodeCodex
Design and scaffold a new trapstreet.run task to evaluate a given agent/skill/tool -- the reverse of trapstreet-solution-scaffold (a solution for an existing task). Generates the mechanical parts (traptask.yaml, judge.py/grader.py on the TRAPTASKMANIFEST contract, buildcases.py's validate-then-render pipeline) and…
Skill Claude CodeCodex
Use when verifying that an implementation matches a change's SDD artifacts. Triggers: "verify", "check implementation", "did I implement everything", "verify the change", "is implementation complete", "check conformance".
Skill Codex
Translate natural-language browser automation requests into exact playwright-cli commands for interactive web testing and debugging. Use when requests involve opening/navigating pages, interacting with elements, capturing snapshots/screenshots/PDFs, using tabs, inspecting console/network, mocking routes, managing…
colinwilliams91/gaitor-orchestrator-cli
Skill Claude Code
Automate browser interactions, test web pages and work with Playwright tests.
Skill Claude Code
Browser-based UI verification using Playwright. Page Object Model, selector best practices, visual regression, network interception, and MCP integration. Trigger: When writing E2E tests, verifying UI changes, or setting up Playwright.
Skill Claude CodeCodex
A debugging workflow that reproduces a problem with a failing test, collects evidence, finds the underlying cause, applies a fix, and runs regression checks. Regression checks verify that the fix does not break behavior that already worked.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: