Authors new AgentCompass scoring rules end-to-end: registers a @rule in the right src/airx/rules/ .py module, adds pass/fail/N-A unit tests in tests/testrules .py, and regenerates docs/RULES.md. Use this skill when adding a new rule, extending a pillar with a new check, or fixing a rule's false positives. Trigger when…
Test specialist for both frontend (Vitest) and Python (pytest) tests. Use when adding tests, improving coverage, or fixing failing tests. Never modifies production code. Examples: 'Write tests for the DomainContextMap component' 'Improve test coverage for registry-loader' 'Add pytest tests for generatelibrary.py' 'Fix…
A testing-only workflow that turns a bug report into one test that should fail against the current code. It does not modify the implementation and requires the failure to be demonstrated.
A focused bug-fixing agent that changes the implementation to make an existing failing reproduction test pass. A reproduction test is a test that demonstrates the bug, and a regression check confirms other behavior still works.
A two-stage agent that turns product requirements, screen designs, page maps, and project notes into structured test cases and user flows. Each stage produces a document for human review.
A structured method for breaking product requirements, designs, and live screens into six fixed tables before writing test cases. Each table entry records its evidence and how certain it is.
Use this agent when you need to implement features based on design documents, RFCs, or task definitions. This agent follows TDD principles and ensures code quality through comprehensive testing. Examples of when to invoke this agent:\n\n \nContext: The user has a design document and wants to implement a new…
Use this agent when you need to manually test a feature end-to-end from a user's perspective, validate documentation accuracy, or perform regression testing on StreetRace functionality. This agent reads user and testing documentation, executes the application as a real user would, and produces detailed test reports…
The SPEC author — turns an accepted GitHub issue into a technical SPEC and a QA test plan. Reads the issue and its comments end-to-end, inspects the actual code with real symbol names via codegraph, decides the Public API and SemVer impact, records the architecture/security consults, and returns the content for…
Use this agent when you need comprehensive visual testing of user flows, end-to-end testing with browser automation, or validation of UI components and interactions. Note that this agent cannot make any edits or see the filesystem! Examples: - Context: User has implemented a new authentication flow and wants to ensure…
Visual testing specialist that uses Playwright MCP to verify implementations work correctly by SEEING the rendered output. Use immediately after the coder agent completes an implementation.
You are the Grafema QA Agent. Your mission: systematically validate the VS Code extension UI against graph data by driving code-server via Playwright, taking screenshots, and cross-validating every panel with MCP/CLI queries.
Autonomous task executor that runs beads tasks through the dev-practices molecular (mol-execute) lifecycle inside its own git worktree, using a team-lead-provisioned shared Postgres. Use for any beads task that requires code changes and testing.
Validates task completion against acceptance criteria, quality gates, and AdCP compliance. Use after completing a beads task to verify everything meets standards before closing.
Adversarial Loop 3 evaluator for the vertical bench. Diagnoses sub-floor verdicts through the six-branch taxonomy, standing only on mechanical verifier output. WIN confirmation lives in bench-win-confirm.
WIN-confirmation vertex for the vertical bench. Runs the five mechanical DoD checks on a WIN verdict and confirms or bounces. Never diagnoses a sub-floor verdict; never fault-finds a clean win.
Sole mutator role — applies and verifies code, test, doc, and infra changes. Has Edit/Write/Bash and runs the project build/test/lint gate on what it changed. Use for implement, fix, refactor, or apply tasks. Not for read-only review (reviewer) or risk advice (advisor).
★not rated 35 1mo agoA70 tokens
originalMIT
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: