test-macos-app
217Command ✓ vendor
Run the smallest meaningful macOS test scope first and explain failures by category.
28,585 tagged Testing, measured the same way as everything else here.
Browse within: accessibility 13Evaluation 12ai-testing 10qa 10Structured Output 9api-testing 9browser-automation 9token-efficiency 9AI Safety 7a11y 7ci 7code-quality 7regression-testing 7agent-evaluation 6
Command ✓ vendor
Run the smallest meaningful macOS test scope first and explain failures by category.
Skill Claude CodeCodex
Scaffolds eval.yaml evaluation specs for agent skills in the dotnet/skills repository. Use when creating skill tests, writing evaluation stimuli, defining graders and rubrics, sizing an eval for statistical power, or setting up test fixture files. Handles the Vally eval.yaml schema, fixture organization, and…
Command
Enforce test-driven development workflow. Scaffold interfaces, generate tests FIRST, then implement minimal code to pass. Ensure 80%+ coverage.
Command
Make fraimz for a flow — run the eval loop and output frame-by-frame proof (fraimz.html).
Agent
Specialized in writing comprehensive test suites. Use for creating unit tests, integration tests, and test documentation.
Skill Claude CodeCodex
Scores how completely an implementation fulfills a PRD/spec, case by case, and produces a single comparable final grade. Invoke only when explicitly named (e.g. run spec-driven-eval); do not auto-trigger. Use when benchmarking spec-driven implementations, grading acceptance criteria, evaluating whether a feature was…
Skill Claude CodeCodex
Use the Crabbox wrapper for OpenClaw remote validation across Linux, macOS, Windows, and WSL2, including delegated Blacksmith Testbox proof. Report the actual provider and id.
Agent Claude Code
TDD Developer agent - implements features using test-driven development and clean code principles.
Command Claude Code
TDD Developer agent - implements features using test-driven development and clean code principles.
Plugin Claude Code
E2E test triage, debugging, and fix implementation toolkit.
Skill Claude CodeCodex
Run all three agent integration phases sequentially: research, write-tests, and implement using E2E-first TDD (unit tests written last). For individual phases, use /agent-integration:research, /agent-integration:write-tests, or /agent-integration:implement. Use when the user says "integrate agent", "add agent…
Skill Claude CodeCodex
Use this skill to test strategy changes against a fresh test repository. Invoke when the user asks to "test against a test repo", "validate the changes", or wants to verify session hooks, commits, and checkpoint creation work correctly.
Skill Claude CodeCodex
Run the Kiln pre-release smoke test suite plus the standard CI checks (checks.sh), diagnose every prerelease test that broke (and why), and write a clean readable report with recommended actions. Read-only — it never edits code. Use when the user wants to validate a release candidate, run prerelease tests, or asks for…
Skill Claude CodeCodex
Records iOS/Android native MP4 evidence for test/repro flows extracted from an Expensify GitHub PR or issue. Use when the user asks to "record the flow for PR.
Skill Claude CodeCodex
Drive iOS and Android devices for the Expensify App - testing, debugging, performance profiling, bug reproduction, and feature verification. Use when the developer needs to interact with the mobile app on a device.
Command Claude Code
Run tests for enter.pollinations.ai service.
Instructions file
Instructions for PeonPing/peon-ping, covering claude.md, commands, run all tests (requires bats-core: brew install bats-core), run a single test file and run a specific test by name.
mock-server/mockserver-monorepo
MCP server Claude CodeCodexCursor +2
MCP server "mockserver" as configured in mock-server/mockserver-monorepo. Runs in Docker (docker.io/mockserver/mockserver:7.6.0).
Skill Claude CodeCodex
Run and manage the Zep eval harness pipeline — document chunking, user ingestion, document ingestion, evaluation, graph inspection, and results analysis. Use when the user asks to run eval harness scripts, use the Zep eval harness, get terminal commands for eval harness operations, chunk documents, ingest users or…
Instructions file CodexOpenCode
AGENTS.md instructions for harbor-framework/harbor, covering claude.md - harbor framework, contributing, project overview, quick start commands and install.
Skill Claude CodeCodex
Scaffold a new Harbor benchmark adapter by running harbor adapter init and then guide implementation using the Adapters Agent Guide as the authoritative spec.
Skill Claude CodeCodex
Write Harbor task verifiers using Reward Kit. Use when creating or editing a task's tests/ directory, adding grading criteria, setting up LLM/agent judges, or designing verifiers that produce a reward score.
Instructions file CodexOpenCode
AGENTS.md instructions for ThePrimeagen/99, covering testing and e2e / integration style testing.
Skill Claude CodeCodex
Run the agent release gate — a portable, wire-level QA harness for the agent runtime. Drives the same product endpoint the playground drives and asserts on the SSE frame stream and real side effects, never on model prose, so it works against any deployment (cloud or self-hosted) from three env vars. Use before an…
At most 3 mods per repository are shown here — the rest are on their repository pages: