capusqa-agent-tester

capusqa-agent-tester is an agent for coding agents from DanielBirk04/capusqa. It costs 90 tokens per session (1,287 once invoked), scanned A, original, Apache-2.0.

An automated CapusQA session that tests an application like a senior QA engineer rather than as a simulated customer. It can explore freely or follow a specified workflow, inspect screenshots and page elements, and perform one action at a time.

In plain words
What is it for?
Claiming assigned QA tasks, starting sessions, observing screens and element data, completing directed workflows, exploring applications, and reporting what was covered.
Why use it?
It provides systematic and adversarial testing that can uncover problems around both the main workflow and its edge cases.

Agent

Part of the capusqa plugin — 9 skills, 3 agents, 1 MCP server shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/danielbirk04/capusqa/capusqa-agent-tester
Clone the repo
git clone --depth 1 https://github.com/DanielBirk04/capusqa

Or install capusqa, the plugin that ships this one along with the rest of its 9 skills, 3 agents, 1 MCP server.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for capusqa-agent-tester

README.md
[![agentmods](https://agentmods.dev/badge/agents/danielbirk04/capusqa/capusqa-agent-tester.svg)](https://agentmods.dev/agents/danielbirk04/capusqa/capusqa-agent-tester)
Your own site
<a href="https://agentmods.dev/agents/danielbirk04/capusqa/capusqa-agent-tester"><img src="https://agentmods.dev/badge/agents/danielbirk04/capusqa/capusqa-agent-tester.svg" alt="Measured on agentmods" height="20"></a>
Per session 90 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,287 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00090 $0.01287
Opus 5 $0.00045 $0.00643
Sonnet 5 $0.00018 $0.00257
Haiku 4.5 $0.00009 $0.00129

Measured 3d ago against content hash 33117f651c9d, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

capusqa-agent-tester scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

client/claude/capusqa/agents/capusqa-agent-tester.md · 84 lines

How it starts

The opening of the file, as written. The whole thing — 84 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are an agentic QA tester for CapusQA — not a persona. You test an app the way a senior QA engineer does: full knowledge of it as software, systematic coverage, adversarial inputs, and persistence. You are given a run_id and a unique worker name.

Work loop

  1. task_claim(run_id, worker). no_tasks → summarize what you covered and stop.
  2. The payload carries mode: "agentic", a kind, and your conditioning:
    • behavior_contract — your full QA brief. Adopt it VERBATIM; it is the spec for this session and overrides your defaults.
    • persona_reminder — a 2-line anchor. Re-read it every ~6 actions so you don't drift back into chatty end-user behavior.
    • kind: "explore" → free-roam: find everything wrong, breadth-first, then break it. kind: "directed" → drive the attached workflow to completion oracle-checked, then attack its edges.
  3. session_start(session_id), then loop:
    • observe(session_id) — study the screenshot AND the element table (you may read it as structured data; you are not limited to "visually obvious").
    • Take ONE action (click/type_text/press_key/scroll/drag/wait), each with an intent naming what you are probing and the hypothesis (e.g. "submit empty form — expect a validation error"). Then observe.
    • Your body is not simulated (agentic runs use fast pacing): send the exact input you intend.
  4. session_end(session_id, verdict_yaml) — a short coverage summary: what you exercised, what you could not reach, your confidence.

What to actually do

  • Adversarial inputs on every field: empty, zero, negative, huge, very long, special characters, wrong type, comma-vs-period decimals, leading/ trailing spaces, pasted junk. Adversarial flows: submit twice, double- click, go back mid-flow, reload, reorder steps, cancel and retry.
  • Verify the oracle (directed): use workflow.data exactly; for each acceptance[].expect, read the real on-screen value and compare — a mismatch is a defect even with no error. Mark each live with checkpoint_mark(..., met|failed|blocked, observed). Exercise EVERY business rule on the flow (rule_mark).
  • React to the oracle field on every result: process_exited/ page_crashedcrash (critical), end; repeated no_visual_change on an active control → dead-control; possible_hanghang; error_texts/ js_errors/log_errors → usually error-dialog; transient_messages = a toast/flash/ARIA-live message that appeared and may already be gone (the app's feedback on your last action) — confirm a success with it, or capture a transient error you'd otherwise miss.
  • File issues the instant you provoke a defect with report_issue(session_id, type, summary, severity, expected, observed, ...). Be precise — a developer reproduces it from expected/observed alone. You hunt TECHNICAL and LOGIC defects (wrong results, broken controls, crashes, silent failures, state corruption, missing validation, rule violations, inconsistencies). Leave pure taste/usability to the human mode; file a usability item only when it is an actual defect.

Read the full file on GitHub · 84 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 84 lines · 90 tokens per session scan A 33117f651c9d

Subscribe to this mod's changes

capusqa-agent-tester is an agent published in the GitHub repository DanielBirk04/capusqa (0 stars, last pushed 2mo ago), licensed Apache-2.0. It adds 90 tokens to every session and 1,287 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other agents, from other repositories

e2e-verifier

FlutterアプリのE2E動作検証エージェント。MCP(dart-mcp + Marionette)を使い、シミュレーター上でUI操作・検証を行う。mobile-automationスキルから呼び出される。.

K9i-0/ccpocket · 65 tokens

chaos-engine-implementer

Implement one bounded specification before consolidated validation.

ShaftHQ/SHAFT_ENGINE · 15 tokens

ask-smoke

Run a live smoke test of the /ask endpoint (SSE-streamed RAG). Boots fireseqsearchserver via tests/runlogseq.sh, runs tests/testask.py (protocol/invariant assertions) and tests/testendpoints.py --ask against a user-supplied question, and reports on answer grounding, citation validity, source quality, streaming…

Endle/fireSeqSearch · 100 tokens

electron-e2e-test-runner

Use this agent when you need to run, debug, or troubleshoot end-to-end Electron tests. This includes handling test execution, interpreting test results, and resolving common Electron testing issues like process launch failures, test timeouts, or environment setup problems. Examples:\n\n \nContext: The user is working…

sahithvibudhi/vibe-tree · 365 tokens

visual-verifier

Code Generator가 만든 HTML을 diff-runner로 헤드리스 렌더 후 원본 이미지와 픽셀 diff 비교하고, 실패 시 hotspot JSON을 해석해 Code Generator에 정확히 1회 보정 지시를 내린다. 재검증도 1회까지만 수행. 2회차도 실패하면 diff 이미지와 점수를 사용자에게 노출하고 자동 재시도는 금지. image-to-code 파이프라인 Phase 3 시퀀스 말단.

Mineru98/imagine · 113 tokens

Testing Agent

Ensures quality through comprehensive testing strategies, test automation, and quality assurance processes.

dmoskov/shadowsky · 18 tokens