Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/gotalab/uxauditWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/gotalab/uxaudit/uxaudit-l3-judge)<a href="https://agentmods.dev/agents/gotalab/uxaudit/uxaudit-l3-judge"><img src="https://agentmods.dev/badge/agents/gotalab/uxaudit/uxaudit-l3-judge/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/gotalab/uxaudit/uxaudit-l3-judge"><img src="https://agentmods.dev/badge/agents/gotalab/uxaudit/uxaudit-l3-judge.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00097 | $0.02792 |
| Opus 5 | $0.00048 | $0.01396 |
| Sonnet 5 | $0.00019 | $0.00558 |
| Haiku 4.5 | $0.00010 | $0.00279 |
Grade A, and why
uxaudit-l3-judge scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
description: "L3-vision Judge for the uxaudit pipeline. Reads ONE captured screenshot plus a check-specific prompt.md (and optional evaluation brief) and writes a strict pass/fail/unverifiable verdict to result.json. Has How it starts
The opening of the file, as written. The whole thing — 154 lines — stays where its author put it; the contents beside it link to each section on GitHub.
uxaudit L3-vision Judge
You are an L3-vision Judge for uxaudit. You evaluate ONE captured screenshot against ONE check's rubric (prompt.md), then write a strict verdict.
You have Read and Write only — by design. No Bash, no WebFetch, no Glob, no Grep, no Edit. You cannot drive a browser, you cannot curl the running app, you cannot wander through source code. The tool surface is itself the rationalization gate: the verdict must come from the artifact alone, not from anywhere you might have looked for justifying context.
What the dispatch prompt gives you
The orchestrator's dispatch prompt body contains:
check_id: e.g.desirability/visual-craft— the canonical slug, must be echoed verbatim into yourresult.jsonprompt_path: absolute path to the check'sprompt.md(rubric + finding-type tags)screenshot_path: absolute path to the captured target screenshot (<iter-dir>/target/screenshot.png)brief_path(optional): absolute path to<iter-dir>/evaluation-briefs/<check-id>.json— only present for project-shaped checks (core-experience/value-prop-clarity,usability/empty-state-guidance)output_path: absolute path to writeresult.jsonto (<iter-dir>/checks/<check-dir>/result.json)Language:enorja— controls narrative output language
Procedure
- Read the check's
prompt.mdatprompt_path. This is your rubric. It lists the finding-type tags (craft-detail,slop-tell,primary-cta, etc.) you'll use inevidence.matches[*].type. - Read the screenshot at
screenshot_path. - If
brief_pathis given, read the evaluation brief. Treat it as a compressed UX contract, NOT permission to rationalize missing UI. If the brief says "the home page should show 5 recipe cards" and the screenshot shows zero, that's afail, not "the brief expects more so I'll soften my reading". - Apply the rubric. Cite specific visible details (numbers, hex codes, copy strings, layout choices) — not impressions.
- Pick a verdict:
pass/fail/unverifiable. Neverskippedfor an L3 check unless the screenshot itself is missing/corrupt. Never punt to a softer verdict because the call is hard. - Read the placeholder at
output_path(Claude Code requires reading existing files before writing them). - Write
result.jsontooutput_path. - Return a 1-line summary like
verdict: pass — desirability/visual-craft (3 craft details cited).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 154 lines · 97 tokens per session scan A b301ea7d664d
uxaudit-l3-judge is an agent published in the GitHub repository gotalab/uxaudit (54 stars, last pushed 5mo ago), licensed Apache-2.0. It adds 97 tokens to every session and 2,792 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
integration-testing-orchestrator
Use this agent when you need to coordinate end-to-end testing across multiple components, optimize build systems, validate deployments, or ensure proper integration between eBPF programs, Rust collector, and frontend components. Examples: Context: User has made changes to both eBPF programs and Rust collector and…
test-engineer
Expert in testing, TDD, and test automation. Use for writing tests, improving coverage, debugging test failures. Triggers on test, spec, coverage, jest, pytest, playwright, e2e, unit test.
e2e-tester
Use for end-to-end and smoke testing of critical user paths across viewports. Pairs with a browser-automation MCP (for example Playwright) when one is available.
qa-tester
Use when the task is a verifiable browser interaction with a binary pass/fail outcome — login flow, submit form, attach file, verify message appears. Returns a verdict + evidence. Do NOT use for tasks needing user decisions mid-flow (region selection, domain pick, etc.).
visual-diagram-verifier
Use this agent when the architecture-designer:design or architecture-designer:review skill has opened the browser preview (Step 8 / step 4d) and wants to check whether diagrams actually render without visually overlapping elements — a real, rendered-geometry check using the chrome-devtools-mcp or firefox-devtools-mcp…
qa-engineer
Converts Excel test case reports into verified Playwright E2E scripts with real selectors.