uxaudit-l3-judge

uxaudit-l3-judge is an agent for Claude Code from gotalab/uxaudit. It costs 97 tokens per session (2,792 once invoked), scanned A, original, Apache-2.0.

A screenshot-based judge for the uxaudit testing process. It compares one captured screen with a written check and records whether the screen passes, fails, or cannot be judged.

In plain words
What is it for?
Use it to assess individual UX checks after browser tests, such as whether a page communicates its purpose clearly or guides users when there is no content.
Why use it?
It keeps the verdict focused on the visible evidence instead of letting the evaluator inspect code, browse the app, or search for supporting information. uxaudit is a testing system for finding usability, accessibility, journey, and interface-quality problems in web apps and webviews.

Agent for Claude Code

Written for Claude Code: shipped in a Claude Code plugin. Also seen: model in frontmatter; mentions Claude Code.

Part of the uxaudit plugin — 1 skill, 7 agents, 1 hook shipped together

Good fit Use it to assess individual UX checks after browser tests, such as whether a page communicates its purpose clearly or guides users when there is no content.

Compare 6 agents from other repositories ↓
Install with agentmods
npx agentmods add agents/gotalab/uxaudit/uxaudit-l3-judge
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Clone the repo
git clone --depth 1 https://github.com/gotalab/uxaudit

Made for: Claude Code.

Or install uxaudit, the plugin that ships this one along with the rest of its 1 skill, 7 agents, 1 hook.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for uxaudit-l3-judge

README.md
[![agentmods](https://agentmods.dev/badge/agents/gotalab/uxaudit/uxaudit-l3-judge/github.svg)](https://agentmods.dev/agents/gotalab/uxaudit/uxaudit-l3-judge)
Your own site
<a href="https://agentmods.dev/agents/gotalab/uxaudit/uxaudit-l3-judge"><img src="https://agentmods.dev/badge/agents/gotalab/uxaudit/uxaudit-l3-judge/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for uxaudit-l3-judge

Your own site · 80×15
<a href="https://agentmods.dev/agents/gotalab/uxaudit/uxaudit-l3-judge"><img src="https://agentmods.dev/badge/agents/gotalab/uxaudit/uxaudit-l3-judge.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 97 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 2,792 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00097 $0.02792
Opus 5 $0.00048 $0.01396
Sonnet 5 $0.00019 $0.00558
Haiku 4.5 $0.00010 $0.00279

Measured 11d ago against content hash b301ea7d664d, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-11, from the pricing page.

Security

Grade A, and why

uxaudit-l3-judge scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

description: "L3-vision Judge for the uxaudit pipeline. Reads ONE captured screenshot plus a check-specific prompt.md (and optional evaluation brief) and writes a strict pass/fail/unverifiable verdict to result.json. Has
agents/uxaudit-l3-judge.md · 154 lines

How it starts

The opening of the file, as written. The whole thing — 154 lines — stays where its author put it; the contents beside it link to each section on GitHub.

uxaudit L3-vision Judge

You are an L3-vision Judge for uxaudit. You evaluate ONE captured screenshot against ONE check's rubric (prompt.md), then write a strict verdict.

You have Read and Write only — by design. No Bash, no WebFetch, no Glob, no Grep, no Edit. You cannot drive a browser, you cannot curl the running app, you cannot wander through source code. The tool surface is itself the rationalization gate: the verdict must come from the artifact alone, not from anywhere you might have looked for justifying context.

What the dispatch prompt gives you

The orchestrator's dispatch prompt body contains:

  • check_id: e.g. desirability/visual-craft — the canonical slug, must be echoed verbatim into your result.json
  • prompt_path: absolute path to the check's prompt.md (rubric + finding-type tags)
  • screenshot_path: absolute path to the captured target screenshot (<iter-dir>/target/screenshot.png)
  • brief_path (optional): absolute path to <iter-dir>/evaluation-briefs/<check-id>.json — only present for project-shaped checks (core-experience/value-prop-clarity, usability/empty-state-guidance)
  • output_path: absolute path to write result.json to (<iter-dir>/checks/<check-dir>/result.json)
  • Language: en or ja — controls narrative output language

Procedure

  1. Read the check's prompt.md at prompt_path. This is your rubric. It lists the finding-type tags (craft-detail, slop-tell, primary-cta, etc.) you'll use in evidence.matches[*].type.
  2. Read the screenshot at screenshot_path.
  3. If brief_path is given, read the evaluation brief. Treat it as a compressed UX contract, NOT permission to rationalize missing UI. If the brief says "the home page should show 5 recipe cards" and the screenshot shows zero, that's a fail, not "the brief expects more so I'll soften my reading".
  4. Apply the rubric. Cite specific visible details (numbers, hex codes, copy strings, layout choices) — not impressions.
  5. Pick a verdict: pass / fail / unverifiable. Never skipped for an L3 check unless the screenshot itself is missing/corrupt. Never punt to a softer verdict because the call is hard.
  6. Read the placeholder at output_path (Claude Code requires reading existing files before writing them).
  7. Write result.json to output_path.
  8. Return a 1-line summary like verdict: pass — desirability/visual-craft (3 craft details cited).

Read the full file on GitHub · 154 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 11d ago First seen · 154 lines · 97 tokens per session scan A b301ea7d664d

Subscribe to this mod's changes

uxaudit-l3-judge is an agent published in the GitHub repository gotalab/uxaudit (54 stars, last pushed 5mo ago), licensed Apache-2.0. It adds 97 tokens to every session and 2,792 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

integration-testing-orchestrator

Use this agent when you need to coordinate end-to-end testing across multiple components, optimize build systems, validate deployments, or ensure proper integration between eBPF programs, Rust collector, and frontend components. Examples: Context: User has made changes to both eBPF programs and Rust collector and…

eunomia-bpf/agentsight · 0 tokens

test-engineer

Expert in testing, TDD, and test automation. Use for writing tests, improving coverage, debugging test failures. Triggers on test, spec, coverage, jest, pytest, playwright, e2e, unit test.

ashrafmusa/agenticana · 49 tokens

e2e-tester

Use for end-to-end and smoke testing of critical user paths across viewports. Pairs with a browser-automation MCP (for example Playwright) when one is available.

mnzralee/claude-multi-agent-architecture · 41 tokens

qa-tester

Use when the task is a verifiable browser interaction with a binary pass/fail outcome — login flow, submit form, attach file, verify message appears. Returns a verdict + evidence. Do NOT use for tasks needing user decisions mid-flow (region selection, domain pick, etc.).

DevZonayed/Mochi · 60 tokens

visual-diagram-verifier

Use this agent when the architecture-designer:design or architecture-designer:review skill has opened the browser preview (Step 8 / step 4d) and wants to check whether diagrams actually render without visually overlapping elements — a real, rendered-geometry check using the chrome-devtools-mcp or firefox-devtools-mcp…

sembraniteam/claude-plugins · 112 tokens

qa-engineer

Converts Excel test case reports into verified Playwright E2E scripts with real selectors.

odarino/haren · 21 tokens