grade
505Command
Part of polygraph
Grade an MCP server A–F with the open polygraph litmus (runs the harness).
2,999 tagged Testing, measured the same way as everything else here.
Browse within: agentic-workflow 44claude-plugin 33ai-development 32agentic-coding 31code-quality 31spec-driven-development 27agent-orchestration 26agent-framework 25ai-assistant 25agentic 21ai-workflow 21Multi-Agent 20claude-code-skills 20documentation 20
Command
Part of polygraph
Grade an MCP server A–F with the open polygraph litmus (runs the harness).
Infopibe/everything-claude-code
Command
Part of everything-claude-code
Generate and run E2E tests with Playwright.
Infopibe/everything-claude-code
Command
Part of everything-claude-code
Go TDD workflow with table-driven tests.
Command Cursor
Run comprehensive tests on all CRUX features via LLM interaction and produce a markdown test report.
Command
Generate comprehensive tests including unit tests, table-driven tests, benchmarks, and examples with high coverage.
Command
Generate comprehensive tests for Sinatra routes, middleware, and helpers using RSpec or Minitest.
Harddiikk/dograh-voicelink-plugin
Command
Part of dograh-voicelink-plugin
Verify a VoiceLink ↔ Dograh install — API health, the single WSS endpoint upgrade probe, and (with a token) that voicelink appears in the telephony providers metadata. Read-only.
Command
Part of gh-cli-search
Execute comprehensive test suite for all gh CLI search skills.
Command
Enforce TDD workflow with Pest 4 for Laravel 13.
Command Cursor
Walk every feature-catalog row against routes, roles, and happy paths; report compliance gaps and backlog suggestions.
Command
Lightweight HTTP load test with Locust. Generates scenario, runs for the duration, outputs markdown report with p50/p95/p99 and error rate.
Command
Part of software-factory
Reproduce then TDD-fix a ready-for-agent bug ticket in this checkout. Web bugs get a browser repro first.
Command
Part of coco
Execute the next available tracked task with TDD, pre-commit validation, PR creation, AI code review, and issue tracker bridge sync.
Command
Part of coco
Autonomous execution loop. Runs the TDD cycle repeatedly until all tasks in an epic are complete, with circuit breaker protection and PR workflow.
Command Claude Code
Comprehensive testing setup for FastAPI applications with pytest and async support.
Command
Implement bug fixes, refactors, and features with research, Red-Green-Refactor, docs, verification, security review, and code review.
peterblazejewicz/claude-plugins
Command
Part of dotnet-skills
Implement tasks incrementally in .NET — RED/GREEN/REFACTOR, dotnet build, dotnet test, commit. Add "auto" to run the whole plan in one approved pass.
peterblazejewicz/claude-plugins
Command
Part of dotnet-skills
List the .NET skills, subagents, and lifecycle commands shipped by this plugin (adapted from addyosmani/agent-skills — spec-driven development, TDD, code review, security audit, test engineering, and more with .NET 8+, C# 12+, xUnit/MSTest, EF Core, Avalonia framing).
DanaSandman/claude-code-plugins
Command
Part of accessibility-plugin
Quick accessibility check on a specific file or directory.
Consiliency/Flutter-Structurizr
Command Claude Code
Provide a comprehensive testing report after completion, including.
zhouziyue233/great-econometrics
Command
Part of great-econometrics
Phase 8 Robustness, Heterogeneity & Mechanism Tests. Reads model-spec.md, diagnosticreport.md, and results-memo.md to build a personalised checklist, then runs method-specific robustness checks, heterogeneity analysis, and mechanism tests; generates code; and produces robustness-report.md.
Command Claude Code
Part of memory-mcp
Comprehensive interactive testing of all Memory MCP features.
Command
Design an evaluation for an AI agent or LLM feature: what to test, how to grade it, and how to catch regressions.
Command
Orchestrate an autonomous, defense-in-depth TDD multi-agent team (requirements → spec-sanity gate → planner → coder/reviewer loop → verification/security/validation/judgment gates → human acceptance → retrospective) to build a feature end-to-end against fully verified, independently validated requirements.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: