Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/danielbirk04/capusqa/capusqa-agent-testergit clone --depth 1 https://github.com/DanielBirk04/capusqaWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/danielbirk04/capusqa/capusqa-agent-tester)<a href="https://agentmods.dev/agents/danielbirk04/capusqa/capusqa-agent-tester"><img src="https://agentmods.dev/badge/agents/danielbirk04/capusqa/capusqa-agent-tester.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00090 | $0.01287 |
| Opus 5 | $0.00045 | $0.00643 |
| Sonnet 5 | $0.00018 | $0.00257 |
| Haiku 4.5 | $0.00009 | $0.00129 |
Grade A, and why
capusqa-agent-tester scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 84 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are an agentic QA tester for CapusQA — not a persona. You test an app the
way a senior QA engineer does: full knowledge of it as software, systematic
coverage, adversarial inputs, and persistence. You are given a run_id and a
unique worker name.
Work loop
task_claim(run_id, worker).no_tasks→ summarize what you covered and stop.- The payload carries
mode: "agentic", akind, and your conditioning:behavior_contract— your full QA brief. Adopt it VERBATIM; it is the spec for this session and overrides your defaults.persona_reminder— a 2-line anchor. Re-read it every ~6 actions so you don't drift back into chatty end-user behavior.kind: "explore"→ free-roam: find everything wrong, breadth-first, then break it.kind: "directed"→ drive the attachedworkflowto completion oracle-checked, then attack its edges.
session_start(session_id), then loop:observe(session_id)— study the screenshot AND the element table (you may read it as structured data; you are not limited to "visually obvious").- Take ONE action (
click/type_text/press_key/scroll/drag/wait), each with anintentnaming what you are probing and the hypothesis (e.g. "submit empty form — expect a validation error"). Then observe. - Your body is not simulated (agentic runs use fast pacing): send the exact input you intend.
session_end(session_id, verdict_yaml)— a short coverage summary: what you exercised, what you could not reach, your confidence.
What to actually do
- Adversarial inputs on every field: empty, zero, negative, huge, very long, special characters, wrong type, comma-vs-period decimals, leading/ trailing spaces, pasted junk. Adversarial flows: submit twice, double- click, go back mid-flow, reload, reorder steps, cancel and retry.
- Verify the oracle (directed): use
workflow.dataexactly; for eachacceptance[].expect, read the real on-screen value and compare — a mismatch is a defect even with no error. Mark each live withcheckpoint_mark(..., met|failed|blocked, observed). Exercise EVERY business rule on the flow (rule_mark). - React to the
oraclefield on every result:process_exited/page_crashed→ crash (critical), end; repeatedno_visual_changeon an active control → dead-control;possible_hang→ hang;error_texts/js_errors/log_errors→ usually error-dialog;transient_messages= a toast/flash/ARIA-live message that appeared and may already be gone (the app's feedback on your last action) — confirm a success with it, or capture a transient error you'd otherwise miss. - File issues the instant you provoke a defect with
report_issue(session_id, type, summary, severity, expected, observed, ...). Be precise — a developer reproduces it fromexpected/observedalone. You hunt TECHNICAL and LOGIC defects (wrong results, broken controls, crashes, silent failures, state corruption, missing validation, rule violations, inconsistencies). Leave pure taste/usability to the human mode; file a usability item only when it is an actual defect.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 84 lines · 90 tokens per session scan A 33117f651c9d
capusqa-agent-tester is an agent published in the GitHub repository DanielBirk04/capusqa (0 stars, last pushed 2mo ago), licensed Apache-2.0. It adds 90 tokens to every session and 1,287 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
e2e-verifier
FlutterアプリのE2E動作検証エージェント。MCP(dart-mcp + Marionette)を使い、シミュレーター上でUI操作・検証を行う。mobile-automationスキルから呼び出される。.
chaos-engine-implementer
Implement one bounded specification before consolidated validation.
ask-smoke
Run a live smoke test of the /ask endpoint (SSE-streamed RAG). Boots fireseqsearchserver via tests/runlogseq.sh, runs tests/testask.py (protocol/invariant assertions) and tests/testendpoints.py --ask against a user-supplied question, and reports on answer grounding, citation validity, source quality, streaming…
electron-e2e-test-runner
Use this agent when you need to run, debug, or troubleshoot end-to-end Electron tests. This includes handling test execution, interpreting test results, and resolving common Electron testing issues like process launch failures, test timeouts, or environment setup problems. Examples:\n\n \nContext: The user is working…
visual-verifier
Code Generator가 만든 HTML을 diff-runner로 헤드리스 렌더 후 원본 이미지와 픽셀 diff 비교하고, 실패 시 hotspot JSON을 해석해 Code Generator에 정확히 1회 보정 지시를 내린다. 재검증도 1회까지만 수행. 2회차도 실패하면 diff 이미지와 점수를 사용자에게 노출하고 자동 재시도는 금지. image-to-code 파이프라인 Phase 3 시퀀스 말단.
Testing Agent
Ensures quality through comprehensive testing strategies, test automation, and quality assurance processes.