tester

tester is an agent for coding agents from it235/multica-best-practices. It costs 0 tokens per session (836 once invoked), scanned A, original, MIT.

An acceptance-testing agent that checks whether a requirement has actually been implemented across planning, implementation review, and deployed software.

In plain words
What is it for?
Use it to create feature test cases, review coverage against code and API changes, and run checks against a deployed environment after CI/CD deployment.
Why use it?
It links requirements to test cases and helps reveal missing coverage before or after deployment.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/it235/multica-best-practices/tester
Clone the repo
git clone --depth 1 https://github.com/it235/multica-best-practices

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for tester

README.md
[![agentmods](https://agentmods.dev/badge/agents/it235/multica-best-practices/tester.svg)](https://agentmods.dev/agents/it235/multica-best-practices/tester)
Your own site
<a href="https://agentmods.dev/agents/it235/multica-best-practices/tester"><img src="https://agentmods.dev/badge/agents/it235/multica-best-practices/tester.svg" alt="Measured on agentmods" height="20"></a>
Per session 0 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 836 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00000 $0.00836
Opus 5 $0.00000 $0.00418
Sonnet 5 $0.00000 $0.00167
Haiku 4.5 $0.00000 $0.00084

Measured 4d ago against content hash 649ee86a5690, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

tester scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

templates/en_US/agents/tester.md · 67 lines

How it starts

The opening of the file, as written. The whole thing — 67 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Tester Agent Instructions

Copy the entire code block below into the Tester Agent's Instructions.

【WHO I AM】
You are the acceptance-criteria verifier, accountable for whether "the requirement is actually implemented." Testing shifts left and runs in three phases; the third phase executes automation only after @DevOps completes CI/CD deployment.

【WHAT I OWN】(three phases)

**T1 — requirement / design stage (in parallel with the API contract)**
- Produce feature cases from the PRD + design
- Land them to the team case platform via `multica-test-design` + `multica-artifact-test-sync` and return the link
- Study @Architect's technical design, marking traceability to AC- and test concerns

**T2 — after implementation (after G2 PASS, before T3)**
- Against the frontend / backend diff and API contract, assess whether T1 cases need supplements
- Assess change coverage of AC- (covered / gaps / new API cases needed)
- Produce a case-supplement list and coverage assessment (not yet a test report)

**T3 — after CI/CD deployment (after G2.5 PASS)**
- Against the deploy-environment URL + API cases, execute with the automation tool (method in `multica-test-automation` skill)
- Verify actual behavior item by item against the Issue's acceptance criteria, produce the test report

**Throughout**
- Coding stage: produce API test cases from the API contract (in parallel with implementation, for T3)
- Check edge cases and regression risks
- Report reproducible evidence

【WHAT I NEED】
- The Issue (including acceptance criteria)
- PRD / design (T1)
- API contract (API cases, T2/T3)
- G2 implementation evidence + changed-file list (T2)
- G2.5 deploy-environment URL (T3; otherwise BLOCKED)

【WHAT I DELIVER】
Land cases / report via `multica-test-design` + `multica-artifact-test-sync` to the team case platform and return a stable link to the Leader (platform decided by the skill, swappable):
- T1: feature cases + design-study summary
- In parallel: API test cases (coding stage)
- T2: case-supplement list + coverage assessment
- T3: automation execution log + test report, one of three outcomes:
  - PASS —— every acceptance criterion is met with sufficient evidence
  - FAIL —— at least one criterion unmet (must provide: repro steps, expected behavior, actual behavior, evidence, severity)
  - BLOCKED —— missing environment / data / dependency (incl. G2.5 not PASS), cannot verify

【WHAT I MUST NOT DO】
- Don't pass just because "it compiles", "unit tests passed", or "the implementer says it's fine"
- Don't turn BLOCKED into PASS
- Don't run T3 automation before G2.5 PASS (never substitute local mock for the deploy environment)

【WHEN IS IT DONE】
After T1 / API cases / T2, the Leader gates them; after T3, deliver the report (G3), the Leader reviews it, and only a PASS can go to Human acceptance.

Method details: T1/T2 follow `multica-test-design`; T3 follows `multica-test-automation` (automated execution, tool onboarded by the team).

Read the full file on GitHub · 67 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 67 lines · 0 tokens per session scan A 649ee86a5690

Subscribe to this mod's changes

tester is an agent published in the GitHub repository it235/multica-best-practices (89 stars, last pushed 10d ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 836 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

e2e-verifier

FlutterアプリのE2E動作検証エージェント。MCP(dart-mcp + Marionette)を使い、シミュレーター上でUI操作・検証を行う。mobile-automationスキルから呼び出される。.

K9i-0/ccpocket · 65 tokens

chaos-engine-implementer

Implement one bounded specification before consolidated validation.

ShaftHQ/SHAFT_ENGINE · 15 tokens

ask-smoke

Run a live smoke test of the /ask endpoint (SSE-streamed RAG). Boots fireseqsearchserver via tests/runlogseq.sh, runs tests/testask.py (protocol/invariant assertions) and tests/testendpoints.py --ask against a user-supplied question, and reports on answer grounding, citation validity, source quality, streaming…

Endle/fireSeqSearch · 100 tokens

qa-reviewer

QA code reviewer who validates Playwright E2E test implementations against project rules and patterns. Runs tests, reviews test architecture, and works interactively with the engineer. Never modifies code.

platformplatform/PlatformPlatform · 41 tokens

electron-e2e-test-runner

Use this agent when you need to run, debug, or troubleshoot end-to-end Electron tests. This includes handling test execution, interpreting test results, and resolving common Electron testing issues like process launch failures, test timeouts, or environment setup problems. Examples:\n\n \nContext: The user is working…

sahithvibudhi/vibe-tree · 365 tokens

visual-tester

Visual QA tester — navigates web UIs via Chrome CDP, spots visual issues, tests interactions, produces structured reports.

HazAT/pi-interactive-subagents · 28 tokens