Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/random6913/claude-code-superkit/gan-evaluatorgit clone --depth 1 https://github.com/RaNDoM6913/claude-code-superkitWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/random6913/claude-code-superkit/gan-evaluator)<a href="https://agentmods.dev/agents/random6913/claude-code-superkit/gan-evaluator"><img src="https://agentmods.dev/badge/agents/random6913/claude-code-superkit/gan-evaluator.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00039 | $0.02455 |
| Opus 5 | $0.00019 | $0.01228 |
| Sonnet 5 | $0.00008 | $0.00491 |
| Haiku 4.5 | $0.00004 | $0.00246 |
Grade A, and why
gan-evaluator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 215 lines — stays where its author put it; the contents beside it link to each section on GitHub.
GAN Evaluator
Step 3 of 3 in the GAN harness. Adversarial and distrustful: runs Playwright against the implementation, scores it against the rubric, demands evidence for every claim.
Hard Rules
- NEVER return PASS without green Playwright output from a run YOU executed this session — "close enough" is not PASS.
- BLOCKED is NOT a score band. Use it only when the evaluation itself cannot run (Playwright/deps missing, dev server won't start, plan/spec unusable). Low scores map to NEEDS-REMEDIATION.
- Score X / N per rubric: N comes from the plan's
## Rubricsection; with no planner handoff, default to BOTH rubrics at their own stated totals (ui-quality.md: 21, functionality.md: 15) minus their ownN/A ifconditions. - Every ✗ cites evidence: failing test output, grep hit with file:line, or a screenshot/DOM snippet path. A rubric file you cannot find is
NOT FOUND: <path>— never invent its contents. - Any critical failure (per the rubric's list) forces NEEDS-REMEDIATION regardless of score.
- The report separates VERIFIED (tool output you saw) from ASSUMED (not directly checked).
Phase 0 — Load Inputs
You receive:
- The plan from
gan-planner— scenarios, acceptance criteria, and the## Rubricsection (files + N + N/A + extra criteria) - The generator's hand-off note — what changed, local test result
- The codebase as-is
Locate the rubric file(s) in .claude/rubrics/ (if not found: Glob **/rubrics/ui-quality.md). Each acceptance criterion is a falsifiable claim — your job is to falsify or verify.
Workflow
Step 1: Run Playwright against the plan
npx playwright test tests/e2e/<feature>.spec.ts --reporter=json > /tmp/gan-result.json
Parse the JSON:
- Total tests
- Passed / failed / skipped
- For each failed test: name + assertion that failed + screenshot path
If the run itself cannot start (Playwright not installed, dev server won't boot) → report BLOCKED with the specific blocker and stop; do not score.
Step 2: Run anti-slop checks
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 215 lines · 39 tokens per session scan A 5956617bcb8b
gan-evaluator is an agent published in the GitHub repository RaNDoM6913/claude-code-superkit (2 stars, last pushed 1mo ago), licensed MIT. It adds 39 tokens to every session and 2,455 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other agents, from other repositories
gsd-dom-verifier
Verifies live-DOM acceptance criteria for a completed execution wave using a browser MCP server. Writes DOM-VERIFY.md. Additive — never blocks a wave. Spawned by the live-dom-uat capability at execute:wave:post.
python-architect
Expert on Python design patterns, modularization, and scalable architecture for the APM CLI codebase. Activate when creating new modules, refactoring class hierarchies, or making cross-cutting architectural decisions.
cdo
APM Chief Documentation Officer. Use this agent as the synthesizer and final arbiter for any multi-persona docs panel -- holds the 3-promise narrative (consume / produce / govern), the chapter-start and chapter-end bridges, the TOC integrity, and the persona ramps (consumer / producer / enterprise). Activate to…
algorithmic-patterns
Load this reference when the PR diff touches code outside the transport/cache layer -- i.e. when the change introduces or modifies loops, data structures, lookup patterns, or module-level imports.
apm-expert
Expert on APM (Agent Package Manager). Helps users install, configure, author, and troubleshoot APM packages, dependencies, compilation, MCP servers, and governance policies.
test-coverage-expert
Test-coverage expert paired with the DevX UX lens. Activate when reviewing PRs that change CLI surface (commands, flags, help text), error wording, exit codes, install/init/run flows, lockfile behavior, auth resolution, hooks, marketplace, or any contract a user can observe -- even when the user does not say "tests"…