evaluate-tests
265Command Claude Code
Evaluate the test suite and codebase against the 30-point test scorecard.
2,999 tagged Testing, measured the same way as everything else here.
Browse within: agentic-workflow 44claude-plugin 33ai-development 32agentic-coding 31code-quality 31spec-driven-development 27agent-orchestration 26agent-framework 25ai-assistant 25agentic 21ai-workflow 21Multi-Agent 20claude-code-skills 20documentation 20
Command Claude Code
Evaluate the test suite and codebase against the 30-point test scorecard.
adrian-d-hidalgo/nestjs-mcp-server
Command Claude Code
Author the technical SPEC and the QA test plan for an accepted GitHub issue — invokes the specifier agent (codegraph inspection, the Public API and SemVer calls, the SDK-types check, architecture/security consults) and writes both to .project/tasks/issue- /. Local files only; the GitHub issue is never modified.…
Command
Part of mk-qa-master
Generate maintainable pytest tests from a URL or mobile screen via mk-qa-master's analyzer.
Command
Part of mk-qa-master
Run a focused subset of the user's test suite via mk-qa-master, surface failures, and offer next steps.
Command Claude Code
Quality assurance skill for systematic bug investigation, reproduction, and documentation.
Command Claude Code
The full pre-push gate for this repo - format, types, tests, guardrails, and the repo's own CI run locally.
Command Claude Code
Run comprehensive validation of the entire ai-dev-standards codebase. This command validates code quality, runs all tests, and performs end-to-end testing that ensures the application works exactly as a user would experience it.
Command Claude Code
Spawn a team of agents to write per-obligation behavioral tests in parallel. Each agent researches, writes, runs, and fixes their test independently. Leader consolidates verified tests, runs quality gates, and commits.
marshall0524/everythingclaudecode
Command
Part of everything-claude-code
Generate and run E2E tests with Playwright.
marshall0524/everythingclaudecode
Command
Part of everything-claude-code
Go TDD workflow with table-driven tests.
Command
Part of headless
Compare legacy and migrated sites during framework migration.
Command
Part of headless
AI-driven functional/E2E testing of a website.
Command Claude Code
Interactive MCP server testing helper for tools and resources.
MichelKerkmeester/skilled-harness__spec-driven-agent-loops
Command Codex
This is the Codex runtime entry point for the /create-benchmark command. The canonical, authoritative command definition lives in the OpenCode tree at.
Command
Part of ralph-dev
Command "integration-tests" from mylukin/ralph-dev, covering add integration tests for end-to-end workflows, acceptance criteria and notes.
Command
Part of fable-mode
Run the fable-mode eval suite (probes → pairwise judge → report). Costs tokens — runs headless claude many times. Optional arg: a probe id substring to filter.
Command
Part of autonomous-dev
Minimal pipeline for test-fixing tasks.
Command
Part of desplega
Create TDD implementation plans with strict Red-Green-Commit/Rollback cycles.
Command Claude Code
Run the verification harness and explain the results as acceptance criteria.
Command Claude Code
Use the dev screenshot script to iterate on UI changes. This builds the Go server with dev tags, starts it on an isolated port, opens the browser, and captures a screenshot for visual verification.
Command
Part of codingbuddy
Execute the implementation plan with TDD.
Command
Part of pagokit
Send synthetic webhook events to a locally running PagoKit integration (valid signature, invalid signature, replay attempt) to verify the handler responds correctly.
Command Claude Code
Part of cc-spec-driven
Generate Maestro test cases from test spec.
Command Claude Code
Part of subcog
Run automated functional tests for Subcog MCP tools. Execute the test suite to validate all MCP tool functionality including memory CRUD, search, knowledge graph, prompts, templates, and maintenance operations.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: