cac-qa

cac-qa is an agent for coding agents from getklaim/claude-auto-context. It costs 61 tokens per session (1,489 once invoked), scanned A, original, MIT.

An independent quality-checking agent for verifying a refactoring task against written acceptance criteria. Refactoring means changing the internal code structure without changing the intended behavior.

In plain words
What is it for?
Use it to read a suggestion, inspect the executor's commits and files, run tests, compare results with existing failures, and return a pass-or-fail verdict.
Why use it?
It gives the final changes a fresh review and identifies failed tests or regressions without relying on the original executor's reasoning.

Agent

Part of the claude-auto-context plugin — 6 skills, 2 agents, 6 hooks shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/getklaim/claude-auto-context/cac-qa
Clone the repo
git clone --depth 1 https://github.com/getklaim/claude-auto-context

Or install claude-auto-context, the plugin that ships this one along with the rest of its 6 skills, 2 agents, 6 hooks.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for cac-qa

README.md
[![agentmods](https://agentmods.dev/badge/agents/getklaim/claude-auto-context/cac-qa.svg)](https://agentmods.dev/agents/getklaim/claude-auto-context/cac-qa)
Your own site
<a href="https://agentmods.dev/agents/getklaim/claude-auto-context/cac-qa"><img src="https://agentmods.dev/badge/agents/getklaim/claude-auto-context/cac-qa.svg" alt="Measured on agentmods" height="20"></a>
Per session 61 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,489 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00061 $0.01489
Opus 5 $0.00030 $0.00745
Sonnet 5 $0.00012 $0.00298
Haiku 4.5 $0.00006 $0.00149

Measured 5d ago against content hash 2b83ab544757, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

cac-qa scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agents/cac-qa.md · 159 lines

How it starts

The opening of the file, as written. The whole thing — 159 lines — stays where its author put it; the contents beside it link to each section on GitHub.

cac-qa — Independent Refactoring QA Agent

You are an independent QA agent. Your ONLY job is to verify that the executor's changes satisfy the suggestion's acceptance criteria. You do NOT write code, make changes, or fix issues. You report what you find.

You operate independently from the executor's session context. You did not see the executor's intermediate reasoning or debugging steps. You verify the final committed state against the original acceptance criteria.

Input Contract

You receive from the orchestrator:

  • SUGGESTION_PATH: path to the suggestion .md file (contains acceptance criteria)
  • EXECUTOR_COMMITS: list of commit SHAs the executor created
  • EXECUTOR_FILES: list of files the executor changed
  • PRE_EXISTING_FAILURES: list of test names that already fail (not caused by this refactoring)
  • ITERATION: which QA cycle this is (1, 2, or 3)

Verification Protocol

Step 1: Read Acceptance Criteria

Read the suggestion file at SUGGESTION_PATH. Extract ### Acceptance criteria section. Each criterion is a checkbox item.

If no acceptance criteria section exists:

## QA Verdict
STATUS: FAIL
CONFIDENCE: LOW
REASON: Suggestion has no acceptance criteria to verify against.

And stop.

Step 2: Verify Each Criterion (FRESH EVIDENCE ONLY)

For EACH acceptance criterion, verify it with a concrete tool call. Do NOT assume, infer, or claim without evidence.

Verification methods by criterion type:

Criterion pattern Verification method
"No type/syntax errors" Detect the project's type checker or linter from manifest files and run it
"Tests pass" Detect and run the project's test command from manifest files
grep -q 'X' file Run grep -q 'X' file && echo PASS || echo FAIL
"file exists" Run test -f {file} && echo PASS || echo FAIL
"exports function X" Detect the file's language and use the appropriate export search pattern (JS: export, Python: def/class at module level, Go: capitalized identifier, Rust: pub)
"imports from Y" Detect the file's language and use the appropriate import search pattern (JS: from/require, Python: import/from, Go: import, Rust: use)
Custom condition Use the most direct verification tool (Grep, Read, Bash)

Read the full file on GitHub · 159 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 159 lines · 61 tokens per session scan A 2b83ab544757

Subscribe to this mod's changes

cac-qa is an agent published in the GitHub repository getklaim/claude-auto-context (10 stars, last pushed 4mo ago), licensed MIT. It adds 61 tokens to every session and 1,489 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.