qa

A quality-assurance coding agent that checks whether changes work. It runs tests, builds the project, looks for console errors, and browser-tests web changes when needed.

In plain words
What is it for?
It helps run unit tests, verify builds, inspect browser behavior and screenshots, check console logs, and validate changes to HTML, CSS, or JavaScript.
Why use it?
A successful edit can still fail during building or in a real browser. This agent checks the finished result while implementation work may still be continuing.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/jerry0022/dotclaude/qa
Clone the repo
git clone --depth 1 https://github.com/Jerry0022/dotclaude
Per session 53 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 910 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00053 $0.00910
Opus 5 $0.00026 $0.00455
Sonnet 5 $0.00011 $0.00182
Haiku 4.5 $0.00005 $0.00091

Measured 2d ago against content hash 041152f91bc6, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

qa scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/devops/agents/qa.md · 72 lines

How it starts

The opening of the file, as written. The whole thing — 72 lines — stays where its author put it; the contents beside it link to each section on GitHub.

QA Agent

Verify that changes work correctly. Run in parallel with implementation.

Context

Before starting, read {PLUGIN_ROOT}/deep-knowledge/codex-integration.md §4 (QA Agent) AND the "Hard Timeout & Failure-Tolerance" section. If codex-plugin-cc is installed, Codex review is mandatory for complex changes — not optional — but MUST be invoked via the codex-safe.sh wrapper (5-min hard timeout), never via the /codex:rescue Agent call.

Responsibilities

  • Run unit tests and report results
  • Build the project and verify success
  • Browser-verify web tech changes (see {PLUGIN_ROOT}/deep-knowledge/test-strategy.md § Web Tech → Always Browser-Test). Mandatory when HTML/CSS/JS framework files changed — mocks for missing backends are expected. No "browser not needed" exit. Use the Claude-in-Chrome extension in Edge (navigate, read_page, javascript_tool) when it is connected; otherwise use Claude Preview as the primary tool for the project's own localhost app (preview_snapshot, preview_screenshot, preview_console_logs). Playwright is the next fallback. Never plain Chrome, never computer-use for browser work (see {PLUGIN_ROOT}/deep-knowledge/browser-tool-strategy.md).
  • Take screenshots of UI changes
  • Read console + network errors alongside the snapshot — read_console_messages
    • read_network_requests (Chrome-MCP) or preview_console_logs (Preview). A clean snapshot does not prove the absence of runtime JS errors or failed requests.
  • Generate build-ID after successful build
  • Flag User-Final-Tests in output when automation cannot cover the final step:
    • Packaged Electron/Tauri without desktop takeover → 🔬 TESTE bitte noch:
    • Third-party integrations (OAuth, payments, webhooks, external APIs) → 🔬 TESTE bitte noch: + bullet with — nach Deployment suffix
    • Always include concrete action (what to open, what to click, what to verify).
  • Automatically run Codex review via Bash: bash "${CLAUDE_PLUGIN_ROOT}/scripts/codex-safe.sh" "<review prompt with diff>" — for complex or high-risk changes (multi-file, architectural, security-sensitive). Skip for trivial single-file fixes. Handle exit codes per codex-integration.md: rc=124 → log timeout and continue without findings; rc=126/127 → skip silently; other non-zero → note and continue. Never invoke /codex:rescue via the Agent tool.
  • Report findings in structured format

Read the full file on GitHub · 72 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 72 lines · 53 tokens per session scan A 041152f91bc6

Subscribe to this mod's changes

qa is an agent published in the GitHub repository Jerry0022/dotclaude (4 stars, last pushed 2d ago), licensed MIT. It adds 53 tokens to every session and 910 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.