Testing agents

28,585 tagged Testing, measured the same way as everything else here.

Browse within: code-quality 57agent-orchestration 47harness 40spec-driven-development 40agentic-workflow 39Multi-Agent 38playwright 36agentic-coding 32github-copilot 31rtl 31verification 31agentic 29copilot 29context-engineering 27

prest/prest

Skill Claude CodeCodexCursor

Guides writing and reviewing pREST Docker/network integration tests under integration/ so each request is human-readable via step comments or table-driven descriptions. Use when adding or editing integration//test.go, HTTP controller E2E coverage, make test-integration, test-integration-postgres…

4.6k +2 5d ago A 78 tokens original MIT

bat-adhoc

242

homeassistant-ai/ha-mcp

Skill Claude CodeCodex

Run bot acceptance tests to validate MCP tools work correctly from a real AI agent's perspective. Use when testing PRs, detecting regressions, or verifying tool changes end-to-end with Claude/Gemini CLIs.

4.6k +29 today A 48 tokens original MIT

bat-story-eval

243

homeassistant-ai/ha-mcp

Skill Claude CodeCodex

Compare MCP tool behavior between target and baseline versions using pre-built and custom stories with diff-based triage.

4.6k +29 today A 26 tokens original MIT

grader

244

EvoScientist/EvoScientist

Agent

Evaluate expectations against an execution transcript and outputs.

4.6k +28 yesterday A 0 tokens copy · 100% Apache-2.0

evaluate-environments

245

PrimeIntellect-ai/verifiers

Skill Claude CodeCodex

Run and evaluate verifiers tasksets. Set up the necessary config files and observe the runs and their results.

4.6k +7 today A 26 tokens original MIT

python-unit-tests

246

dimensionalOS/dimos

Skill Claude CodeCodex

Use when writing, fixing, or reviewing Python pytest unit tests, fixtures, mocks, or test PR feedback.

4.4k +8 today A 26 tokens

sanity-check

247

rivet-dev/agentos

Skill Claude CodeCodex

Run the deferred AgentOS E2E smoke test from public npm packages. Use when the user asks to sanity check, smoke test, or verify a release works.

4.4k +9 2d ago A 37 tokens original Apache-2.0

browser-smoke-review

248

fallow-rs/fallow

Skill Claude CodeCodex

Use browser automation to review docs pages, preview URLs, rendered output, or web-facing fallow surfaces. Use when the user wants a screenshot-based review, browser smoke test, docs site check, or preview deployment inspection.

4.4k +9 yesterday A 49 tokens original MIT

bash

249

imazen/imageflow

Cursor rule Cursor

Rules for working on bash scripts and their tests.

4.4k 4d ago A 678 tokens AGPL-3.0

codex-qa-tester

250

Waishnav/devspace

Agent

Manual QA profile for browser testing, workflow verification, and regression checks.

4.4k +119 2d ago A 21 tokens original MIT

add-test

251

mixedbread-ai/mgrep

Command Claude Code

Command "add-test" from mixedbread-ai/mgrep, covering add test, arguments and steps.

4.4k +1 4mo ago A 0 tokens original Apache-2.0

contributing

252

crmne/ruby_llm

Skill Claude CodeCodex

Contribute to RubyLLM - set up the repo, run and record specs, add providers or chat options, work on the Rails integration, and edit docs. Use when fixing a bug, building a feature, writing specs, or changing documentation in the RubyLLM codebase.

4.3k +7 yesterday A 60 tokens original MIT

testing

253

callstack/agent-device

Agent

Repository-specific testing traps you cannot learn from the test runner alone. Executable gate ownership lives in scripts/check-affected/ and scripts/gate/.

4.3k +28 today A 0 tokens original MIT

android-emulator

254

callstack/agent-device

Skill Claude CodeCodex

Verify and debug native, React Native, Expo, or Flutter apps on an Android Emulator with agent-device. Use when an agent needs to launch an app, inspect its live UI, tap, type, scroll, validate a code change, collect failure evidence, or reproduce a workflow on an Android virtual device.

4.3k +28 today A 65 tokens original MIT

agent-device

255

callstack/agent-device

MCP server Claude CodeCodexCursor +2

MCP server for mobile app automation: verify, control, and debug iOS, Android, TV, and desktop apps. Runs locally from the agent-device npm package.

4.3k +28 today A tokens not measured original MIT

httprunner CLAUDE.md

256

httprunner/httprunner

Instructions file

Instructions for httprunner/httprunner, covering claude.md, project overview, development commands, building and testing.

4.3k +1 8mo ago A 1,029 tokens original Apache-2.0

seed-ssim-references

257

hao-ai-lab/FastVideo

Skill Claude CodeCodex

Seed HF reference artefacts for a single newly-added SSIM test (pixel .mp4 for runtexttovideosimilaritytest-style tests, or latent .pt for runtexttolatentsimilaritytest-style tests). Runs the test on Modal L40S, downloads the generated artefacts via modal volume get, pauses for the user to verify (visual eyeball for…

4.3k +75 yesterday A 146 tokens original Apache-2.0

feishu-e2e-test

258

m1heng/clawdbot-feishu

Skill Claude CodeCodex

Local E2E debug and test framework for clawd-feishu plugin development. Use when debugging message flow, testing bot responses, verifying Feishu web UI interactions, or performing end-to-end validation of the OpenClaw-Feishu integration during development.

4.2k +1 5mo ago A 58 tokens original MIT

archestra-dev-testing

259

archestra-ai/archestra

Skill Claude CodeCodex

Use when deciding whether a change needs a test and at which level — unit, backend route-level integration, MSW-backed frontend integration, or e2e — or when reviewing tests for the "fluff test" anti-pattern. Start here before archestra-dev-backend-tests or archestra-dev-e2e.

4.2k +4 yesterday A 68 tokens

env-graph-tests

260

dmno-dev/varlock

Cursor rule Cursor

Cursor rule "env-graph-tests" from dmno-dev/varlock, covering rule: general test structure for env-graph, description, rationale, requirements and example.

4.2k 3d ago A 0 tokens original MIT

oss-fuzz

261

apache/tika

Skill Claude CodeCodex ✓ vendor

Run Tika's OSS-Fuzz Jazzer targets locally against a working-tree checkout — build the image, build fuzzers from local source, fuzz a target, run a corpus as a regression pass, reproduce a crash, and add seeds. Use for "fuzz the OneNote parser", "run OneNoteParserFuzzer against these files", "reproduce an OSS-Fuzz…

4.0k +10 yesterday A 90 tokens original Apache-2.0

skill-creator

262

OpenBMB/PilotDeck

Skill Claude CodeCodex

Create new skills, modify and improve existing skills, and measure skill performance. Use when users want to create a skill from scratch, edit, or optimize an existing skill, run evals to test a skill, benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy.

4.0k +4 yesterday A 64 tokens AGPL-3.0

comparator

263

OpenBMB/PilotDeck

Agent

Compare two outputs WITHOUT knowing which skill produced them.

4.0k +4 yesterday A 0 tokens AGPL-3.0

grader

264

OpenBMB/PilotDeck

Agent

Evaluate expectations against an execution transcript and outputs.

4.0k +4 yesterday A 0 tokens AGPL-3.0

At most 3 mods per repository are shown here — the rest are on their repository pages: