antrieb
673MCP server Claude CodeCodexCursor +2
Validates AI infra code on real VMs. Self-corrects until it works. No containers, no sandboxes. Remote server at antrieb.sh.
18,346 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.
MCP server Claude CodeCodexCursor +2
Validates AI infra code on real VMs. Self-corrects until it works. No containers, no sandboxes. Remote server at antrieb.sh.
Skill Claude Code
Run smoke tests against the mpak CLI to verify all commands work correctly before publishing a new release.
Plugin Claude Code
Bundles 5 skills, 2 commands, 1 hook, 1 MCP server · 827 tokens together
Connect to Kobiton mobile testing platform - manage devices, run automation suites, and view test results.
Agent Claude Code
Use this agent when you have created a plan (e.g., implementation plan, architecture design, refactoring strategy, feature specification) and need to validate and iteratively improve it before execution. This agent should be invoked:\n\n- After drafting any significant technical plan that will guide implementation…
Command Claude CodeCursor
This command performs comprehensive code quality checks. Use it before commits or when implementation is complete.
LuthienResearch/luthien_control
Cursor rule Cursor
This rule outlines the strict Test-Driven Development process to be followed by the AI assistant when implementing new features or fixing bugs. This complements the broader developmentworkflow by providing specific TDD execution steps for the AI.
Plugin Claude Code
Bundles 1 skill, 1 hook · 203 tokens together
Deterministic scanner that finds and deletes tests that can never fail — assertions the type checker already guarantees, tautologies, mock-echo tests, dead/swallowed assertions, and duplicates. TypeScript (Jest/Vitest/bun:test) and Python (pytest + mypy).
MCP server Claude CodeCodexCursor +2
HarmonyOS MCP service for device automation, app deployment, UI interaction, E2E support, and log validation. Runs locally from the harmonyos-dev-mcp Python package.
Skill Codex
Audit and prove Bagisto storefront themes with source inspection, ownership mapping, admin-to-storefront mutation tests, and Playwright commerce journeys. Use when checking that visible content is dynamic and merchant-controlled; validating theme customizations, channels, CMS, categories, products, search, filters…
MCP server Claude CodeCodexCursor +2
MCP server for AI-driven Cypress test execution. Run, debug, and iterate on E2E tests directly from your AI agent. Runs locally from the cypress-mcp npm package.
djscheuf/agentic-dev-ecosystem-template
Skill Claude Code needs its repo
Complete E2E test debugging workflow (composite orchestrator). Starts by reviewing the provided test failure evidence, then forms hypotheses, applies fixes, and verifies results for a presumed E2E playwright test suite.
MCP server Claude CodeCodexCursor +2
MCP Server for Swagger/OpenAPI documentation and API testing. Runs locally from the swagger-mcp npm package.
Skill Claude Code
A test-first coding process based on TDD, or test-driven development: write a failing test, make it pass, then improve the code.
Skill Claude Code
Use when adding unit or integration tests. Provides test patterns, naming conventions, and fixtures for Python (pytest), TypeScript (vitest), Java (JUnit/Mockito), and Rust.
Skill Claude Code
Core conventions and rules for Playwright E2E testing with TypeScript. Covers project structure, naming conventions, selector strategy, authentication, navigation, environment configuration, test independence, and parallelism. Automatically loaded when writing or modifying E2E tests. Use when: generating E2E tests…
Command Claude Code ✓ vendor
You are a GEDD coaching assistant. You guide the user through building a golden evaluation dataset for their AI agent using Open Coding methodology, then help them evaluate and annotate responses — all without leaving Claude Code.
Skill Claude CodeCodex
End-to-end workflow for feature requests, enhancements, refactors, and tasks. Covers planning, TDD implementation, verification, integration testing, and PR creation.
Skill Claude Code needs its repo
Autonomous QA tester that systematically tests every rdst command via tmux harness, applies a quality rubric, and files bugs in beads.
Plugin Claude Code
Bundles 38 skills · 617 tokens together
Turn AI software engineering into an auditable, on-disk state machine. A staged PDLC workflow (PRD, design, TDD, implement, review, ship, retro) enforces hard contracts — artifacts on disk, per-feature state machine, tests-before-code, objective checks from real command exit codes, single-shot auto-repair — so AI work.
Plugin Claude Code
Bundles 6 skills, 2 agents · 1,019 tokens together
Four agent skills for Playwright and Cypress end-to-end tests: generate new Playwright coverage with live exploration only on local/disposable or externally isolated approved non-production targets, review existing specs or PR diffs, and debug failed runs.
Skill Claude CodeCodex
Drive MANTA's built Electron app in a real renderer context (Playwright's electron launcher) and assert that key UI surfaces render correctly — no crash, no blank screen, sidebar/chat/terminal present. Load BEFORE marking any frontend/UI task done, and when manta-pr-workflow or manta-handle-reviewer-return verifies a…
Skill Codex
Adversarially verify claims that code, fixes, tests, CI, deployments, logs, or systems are correct, complete, healthy, or safe to merge. Use when asked to prove, verify, validate, confirm, double-check, review the agent's own work, check whether a bug is actually fixed, or decide whether green signals justify a…
Plugin Claude Code
Bundles 11 skills, 39 commands, 31 agents, 10 hooks · 5,597 tokens together
Spec-driven Claude Code workflow with verification pipeline. Hooks, slash commands, and verification subagents.
Plugin Claude Code
Bundles 1 skill, 4 agents · 93 tokens together
Multi-agent collaborative development workflow: Planner → Dev → Test → Learn with automatic bug-fix loops, test layering, and experience accumulation.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: