An end-to-end workflow that turns a product requirements document, or MRD, into planned and tested code. It can clarify requirements, create a more detailed plan, design the technical approach, generate code and tests, and archive the work.
Use when starting VAT work or deciding which VAT sub-skill applies. Router that points at sub-skills for adoption, skill/agent authoring, audit, distribution, RAG, knowledge resources, skill review, and enterprise org admin.
Use when you need lightweight browser QA for a web page, local HTML file, or app: inspect console errors, broken assets, keyboard/focus behavior, viewport readability, and publish evidence-backed findings JSON through a local HTML report viewer.
Show and maintain a discrete-node progress panel while ChatGPT uses its built-in browser for an authorized website acceptance test. Use when the user starts an @Browser acceptance task, asks to test a website with the in-app browser, or asks to show browser test progress. Do not use for ordinary browsing, research, or…
A Traditional Chinese teaching guide for a software lab on golden datasets and prompt regression. A golden dataset is a set of expected examples used to check whether changes still produce acceptable results.
Benchmark and evaluate the quality of any AgentSkill. Use when asked to test, evaluate, benchmark, or assess a skill's effectiveness. Triggers on phrases like 'benchmark this skill', 'evaluate skill quality', 'test this skill', 'how good is this skill', 'skill audit', 'skill assessment'. Works by generating test…
Provision temporary Hetzner Cloud VMs for Ansible role QA, bootstrap OS prerequisites such as Rocky 8 Python 3.9, run verified work, clean up cloud resources, and report sanitized evidence.
Evaluate, benchmark, or test-drive an unfamiliar AI agent skill, tool bundle, prompt workflow, or capability package. Use when the user asks whether a downloaded/shared .skill, SKILL.md, agent workflow, or reusable AI capability actually helps, is worth installing, works as advertised, performs better than asking an…
Full project test coverage. Analyzes the entire codebase, identifies coverage gaps, and writes tests to achieve 100% coverage (minimum 80%). Ensures every source file with exportable logic has at least one test. Use when the user mentions 'alltest', wants comprehensive testing, or asks for full test coverage.
Selects the optimal testing mechanism, designs the minimum viable test, and actively executes the creation of the test assets (e.g., landing pages, scripts) using tool calling. Reads the GTM plan and outputs a live test environment and execution timeline.
Deterministic simulation testing for containerized services. Write Lua scripts to inject chaos (pause, kill, resource deprivation) into Docker containers with reproducible, seeded fault injection. Use when writing chaos experiments, testing service resilience, or debugging distributed systems.
Route the aoa-evals skill family for central proof selection, review, evolution, named results or verdicts, source-linked reports, Eval Forge owner review, and proof lifecycle. Hand repository-local eval selection, application, intake/design, or session-hit classification to aoa-eval. Candidates, readiness checks…
How to write good tests - behavior-focused assertions, the AAA structure, red/green/refactor, small vertical slices, and what to mock versus leave real. Use whenever writing a new test, modifying or fixing an existing test, reviewing someone else's tests, deciding whether something needs a test, choosing what to mock…
A skill for creating, editing, checking, and diagnosing Russian-language Turbo Gherkin `.feature` files for Vanessa Automation, a tool for testing 1C applications. It covers UI and export scenarios without MCP.
A guide for documenting and testing an existing software feature or page by studying how its code already works. It creates requirement, design, and test documents that describe the current behavior rather than an imagined future version.
Trigger: "system-optimierung", "system optimieren", "optimierungslauf", "pruefset fahren", "prüfset fahren", "messlauf", "regeltreue messen". Misst mit kalten Prüfset-Läufen (dein Prüfset, beim ersten Mal aus der Vorlage 10System\Pruefset-Vorlage.md angelegt), ob das System seine eigenen Regeln einhält, leitet aus…
Pair on an implementation via ping-pong TDD — alternating red/green rounds between the agent and the user. Use when the user says "ping-pong" or asks to alternate writing failing tests and making them pass.
Benchmark whether Atlas measurably improves coding-agent behavior, by running every scenario twice — with and without Atlas — in clean contexts and scoring the pair blind.
AI-powered QA test automation — record browser flows, generate test cases from PRDs or Figma designs, execute natural language test scripts, and convert to Playwright + Cucumber BDD tests. Use when asked to "record a test", "create test cases", "execute test script", "automate UI test", "convert to BDD", "generate…
★not rated 2 6mo agoA89 tokens
originalMIT
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: