Command Claude Code
An end-to-end test command for Kaku, a tool that manages programming-agent panes, together with Feishu messages and runtime monitoring. It starts from a simulated Feishu message and checks the full agent-team workflow.
2,999 tagged Testing, measured the same way as everything else here.
Browse within: agentic-workflow 44claude-plugin 33ai-development 32agentic-coding 31code-quality 31spec-driven-development 27agent-orchestration 26agent-framework 25ai-assistant 25agentic 21ai-workflow 21Multi-Agent 20claude-code-skills 20documentation 20
Command Claude Code
An end-to-end test command for Kaku, a tool that manages programming-agent panes, together with Feishu messages and runtime monitoring. It starts from a simulated Feishu message and checks the full agent-team workflow.
Command
Part of rn-dev-agent
Execute a learned Maestro flow ("action") by name with optional -e KEY=VALUE parameters. Looks the flow up via packages/rn-dev-agent-core/dist/learned-actions.js (same inventory as /rn-dev-agent:list-learned-actions), then replays it via cdprunaction — auto-repair-aware orchestration with structured RunRecords (GH.
J0hnG4lt/metabase-flightsql-driver
Command Claude Code
Run end-to-end tests for the Metabase Arrow Flight SQL driver. Starts all services from scratch, runs setup, and validates the dashboard works.
Command Claude Code
An interactive evaluation command for scoring and displaying the answer quality of a knowledge agent. An evaluation checks whether an agent gives useful and correct answers.
Command Claude Code
Sequentially author negatives.py for every case that lacks one, with interactive judgment gates where LLM cannot reliably decide.
trainual/tiptap-collaboration-mcp
Command Claude Code
Smoke test the tiptap-collaboration MCP tools against a live server. Use when user says "smoke test", "test tiptap", or "test mcp tools".
Command Claude Code
Generate comprehensive unit testing infrastructure including test configuration, health checks, regression tests, synthetic transactions, coverage dashboards, and CI/CD integration.
Command
Part of polygraph
Run polygen — draft a contract from a feature description, author a verifiable SAM v2 strict-profile module against it, self-repair against reachable invariant violations, and synthesize a demo/regression trace corpus.
Command
Part of polygraph
Run polyvers — classify a state-machine version change into compatibility lanes, run the gates those lanes require against fleet snapshots (shape round-trip, vocabulary, in-flight stimuli, migration validation, seeded model check), scaffold migrations, and check parent×child version matrices. No API key.
Command
Part of polygraph
Run the Polygraph verification loop — generate N transition-function specs from a source file and replay real traces against them, reporting spec-errors vs code-findings.
Command
Part of tailtest
Run an adversarial pass on $ARGUMENTS -- explicitly try to break the source code.
Command
Part of tailtest
Generate or update tests for $ARGUMENTS.
Command Claude Code
Part of canicode
Run the calibration pipeline for a single fixture, all active fixtures, or resume a failed run.
Command Claude Code
Part of canicode
Review a completed /develop pipeline run and provide a structured QA assessment.
Caspian-Sun/claude-code-workflow
Command Claude Code
You are now acting as a Test Engineer. Generate comprehensive test cases for the specified components/functions.
Command Claude Code
Part of sdlc
Run automated functional tests using the hook-driven test framework. Execute the test suite to validate all project functionality.
afhverjuekki/claude-code-aristotle-plugin
Command
Part of aristotle
Verify algorithm correctness using Aristotle (96.8% VERINA benchmark accuracy).
Command
TDD-driven bugfix workflow: tester writes failing test (RED) → developer fixes (GREEN) → developer refactors (REFACTOR) → reviewer validates. Accepts issue number, description, or both. Auto-creates PR unless --no-pr flag is passed.
ai-plugin-marketplace/template
Command
Part of skill-evaluator
Evaluate a skill across model tiers using blind testing.
Command Claude Code
Generate a TDD plan with verbose planning docs and minimal agent execution files using MCP planning tools.
Command Cursor
Command "create-python-tests" from codeready-toolchain/tarsy, covering writing tests for python llm service, running tests, from project root, from llm-service/ directory and critical rules.
Command Claude Code
Part of frontend-expert
Test user-facing UI — frontend-testing, states, a11y; invoke test-engineer.
jesus-seijas-sp/claude-node-setup
Command Claude Code
Run tests for a specific feature or pattern and summarize the results.
Command Claude Code
A command that uses browser automation to run an experience check on the current application or game.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: