llm-testing
241Eyadkelleh/awesome-skills-security
Skill Claude CodeCodex
Comprehensive LLM security testing prompts for bias detection, data leakage, alignment testing, and adversarial prompt resistance.
26,146 mods in this category, of every kind an agent can take. Each one carries what it costs per session, what the scan found, and whether it is the original.
Eyadkelleh/awesome-skills-security
Skill Claude CodeCodex
Comprehensive LLM security testing prompts for bias detection, data leakage, alignment testing, and adversarial prompt resistance.
Instructions file Claude Code
Instructions for tae0y/real-estate-mcp, covering claude.md, project overview, commands, run a single test file and run a single test by name.
currents-dev/playwright-best-practices-skill
Skill Claude CodeCodex
Use when writing Playwright tests, fixing flaky tests, debugging failures, implementing Page Object Model, configuring CI/CD, optimizing performance, mocking APIs, handling authentication or OAuth, testing accessibility (axe-core), file uploads/downloads, date/time mocking, WebSockets, geolocation, permissions…
Skill OpenCode
Dogfood a project-owned tool in a realistic user flow and evaluate the agent experience. Use when the user asks to dogfood or self-iterate, or when testing or evaluating a tool built by this project.
Skill Claude CodeCodex
Automatically generate comprehensive test cases from API definitions, endpoint descriptions, OpenAPI/Swagger specs, Postman collections, or raw HTTP request/response examples. Use this skill whenever the user mentions generating tests from APIs, writing test cases for REST endpoints, API testing, creating test suites…
Command Claude Code
Run a REPL test scenario against the claude-in-mobile REPL plugin (python/node/bash/...).
leeguooooo/claude-code-usage-bar
Instructions file CodexOpenCode
AGENTS.md instructions for leeguooooo/claude-code-usage-bar, covering repository guidelines, project structure & module organization, build, test, and development commands, coding style & naming conventions and testing guidelines.
Skill Claude CodeCodex
A test-design skill for TileLang-Ascend operators, which are custom AI computations for Huawei Ascend chips. It creates or reviews test cases from design documents, example code, user descriptions, or existing tests.
Skill Claude CodeCodex
Automates browser interactions for testing and validating your own web applications using playwright-cli. Use when you need terminal-first browser control for navigation, form filling, screenshots, tracing, bound browser sessions, debugging, or generating Playwright test code. Only use against applications you own or…
Instructions file CodexOpenCode
AGENTS.md instructions for CodSpeedHQ/codspeed, covering agents.md, common development commands, building and testing, build the project and build in release mode.
Skill Claude Code
Guide for adding or updating slime tests and CI wiring. Use when tasks require new test cases, CI registration, test matrix updates, or workflow template changes.
Instructions file CodexOpenCode
Instructions for zksecurity/zkbugs, covering repository guidelines, project structure & module organization, build, test, and development commands, coding style & naming conventions and testing guidelines.
Skill Claude Code
A development guide for adding or improving real-LLM end-to-end tests. End-to-end tests exercise a complete feature path, while a real-LLM test uses an actual language model rather than a fake response.
aminueza/terraform-provider-minio
Instructions file CodexOpenCode
AGENTS.md instructions for aminueza/terraform-provider-minio, covering repository guidelines, quick reference (read first), project structure, build & development commands and run single test.
Skill Claude CodeCodex
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes — find the root cause first instead of guessing at symptom patches.
Skill Codex
Review Benchflow or SkillsBench task-run trajectories and integration-test Benchflow code changes. Use this skill whenever the user asks to audit traj health, failed or timed-out runs, healthy pass/fail/timeout status, no-skill leakage, skill loading, reward hacking, verifier isolation, metadata completeness, token…
Skill Claude CodeCodex
Run, author, or update BlazeDiff visual regression tests. Trigger on "visual test", "screenshot regression", "blazediff", "/blazediff".
Instructions file CodexOpenCode
AGENTS.md instructions for hughjonesd/huxtable, covering instructions to llm agents, documentation, writing code, tests and docs, testing your work and installing extra components.
Instructions file CodexOpenCode
AGENTS.md instructions for adshao/flounder, covering agents.md, project defaults, thin-layer agentic mode, map → dig deep coverage and operating the deep audit.
Instructions file CodexOpenCode
AGENTS.md instructions for smaramwbc/statewave, covering agents.md — guide for contributors and coding agents, what this repo is, setup, build, test, conventions and pull requests.
Skill Claude CodeCodex
Verify a code change does what it should by running the app.
Skill Claude CodeCodex
Test the preview proxy feature end-to-end. Use when verifying preview URL changes, auth changes on the web proxy route, or networking-related fixes.
IncomeStreamSurfer/claude-code-agents-wizard-v2
Agent Claude Code
Visual testing specialist that uses Playwright MCP to verify implementations work correctly by SEEING the rendered output. Use immediately after the coder agent completes an implementation.
Cursor rule Cursor
Cursor rule "style" from Ayanami1314/swe-pruner, covering style guide, bad, good, test style and bad.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: