18,324 mods in this category, of every kind an
agent can take. Each one carries what it costs per session, what the
scan found, and whether it is the original.
Build an LLM-as-Judge evaluator for one specific failure mode. Binary pass/fail only. Use when a failure mode requires interpretation (tone, faithfulness, relevance, completeness) and cannot be checked with code. Do NOT use when the failure can be checked with regex, schema validation, or execution tests. Do NOT use…
Run the checks that have to pass before a FreeReps commit lands — Go build, vet, tests and golangci-lint, the frontend type check and build, and the document contract check. Triggers — "verify", "prüf das durch", "vor dem commit", "läuft das durch", "check before committing", "run the checks", "does CI pass". Not for…
Quadruple verification for Claude Code — automatically blocks placeholder code, security vulnerabilities, and ensures output quality on every operation. Built by CustomGPT.ai for production teams.
★not rated 16▲
+1 5mo agoA
tokens not measured
originalMIT
Autonomously explore the UI of a Visual Studio debuggee via vs-mcp UIA tools (uisnapshot, uifindelements, uiwait) and produce a structured bug/coverage report. Use when the user asks to "test all screens", "crawl the UI", "find UI bugs autonomously", or similar.
Battle-tested Playwright E2E testing patterns for Next.js/React apps. Use when writing, running, debugging, or fixing Playwright tests. Also triggers on 'e2e', 'end-to-end', 'playwright', 'browser test', 'UI test', 'integration test with browser', 'flaky test', 'test keeps failing'. Covers locators, assertions…
Write, edit, review, and validate AgentV EVAL.yaml / .eval.yaml evaluation files. Use when asked to create new eval files, update or fix existing ones, add or remove test cases, configure graders (llm-rubric, script), review whether an eval is correct or complete, convert between EVAL.yaml and evals.json using agentv…
Adversarially verify claims that code, fixes, tests, CI, deployments, logs, or systems are correct, complete, healthy, or safe to merge. Use when asked to prove, verify, validate, confirm, double-check, review the agent's own work, check whether a bug is actually fixed, or decide whether green signals justify a…
AI development loop — orchestrator distributes tasks to headless workers, independent auditor verifies, structural enforcement auto-blocks downstream on upstream failure. Full-cycle validated with 10-scenario test harness.
The Fastest Way to Audit Your RAG - Generate QA datasets & evaluate RAG systems in Colab, Jupyter, or CLI. Privacy-first, any LLM, visual reports. Runs locally from the ragscore Python package. Needs 2 environment variables to run.
★not rated 15 3mo agoA
tokens not measured
originalApache-2.0
Verify CLAUDE.md/AGENTS.md references, compile typed specs, and test the agent harness.
★not rated 15 todayA
tokens not measured
originalMIT
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: