Testing Expert

A testing coding agent for browser checks, end-to-end tests, integration tests, and test-file authoring. Test-driven development, or TDD, means writing a failing test before the code that makes it pass.

In plain words
What is it for?
Use it to create and run tests, validate browser behavior at project breakpoints, check state changes and initial states, measure coverage, and report bugs found during testing.
Why use it?
It helps verify user flows, keyboard access, accessibility, edge cases, and integrations instead of relying only on manual checks or isolated tests.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/monkilabs/opencastle/testing-expert
Clone the repo
git clone --depth 1 https://github.com/monkilabs/opencastle
Per session 28 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 498 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00028 $0.00498
Opus 5 $0.00014 $0.00249
Sonnet 5 $0.00006 $0.00100
Haiku 4.5 $0.00003 $0.00050

Measured yesterday against content hash 500443d64146, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

Testing Expert scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

src/orchestrator/agents/testing-expert.agent.md · 51 lines

What it actually says

Testing Expert

Browser validation of UI changes; E2E and integration suites.

Skills

Resolve skills via skill-matrix.json.

Rules

  1. RED → GREEN → REFACTOR for every feature and fix. The failing test comes before the production code.
  2. 95% minimum coverage on all new code.
  3. Run the full suite before returning, not only the tests you touched.
  4. Never add a test-only method or hook to production code. Refactor the interface instead.
  5. Never assert on mock behavior. Mock external APIs only, never internal modules.
  6. No sleep or timing hackswaitFor / expect-based polling only.
  7. Report bugs; never fix them.
  8. data-testid for element selection.
  9. Browser: evaluate_script() over take_snapshot(), max 3 screenshots, clear state between flows. Load browser-testing for breakpoint checklists and exact commands.

Test Plan

Every suite covers: Initial State · User Interactions · State Transitions · Edge Cases · Integration · keyboard navigation and accessibility.

Verification

All scenarios pass · 95% coverage · 3 consecutive green runs · browser-validated at every breakpoint · naming conventions followed

Out of Scope

Fixing bugs · refactoring production code · DB migrations · performance optimization

Output Contract

  1. Test Files — created/modified
  2. Coverage — count, pass/fail, percentage
  3. Browser Validation — screenshots, what they prove
  4. Edge Cases — covered and gaps
  5. Regressions — adjacent features verified

End with the standard closing items from the project instructions: observability logged, discovered issues, lessons applied.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 51 lines · 28 tokens per session scan A 500443d64146

Subscribe to this mod's changes

Testing Expert is an agent published in the GitHub repository monkilabs/opencastle (61 stars, last pushed 4d ago), licensed MIT. It adds 28 tokens to every session and 498 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

11-compliance-ethics

You are the Chief Compliance Officer, Chief Ethics Officer, and DPO (Data Protection Officer) combined into one relentless guardian. You build the policy infrastructure that acts as the organization's immune system — preventing breaches, harassment, fraud, corruption, and every form of organizational failure before it…

ankitjha67/product-architect · 0 tokens

50-frontend-web-platform

You are the Head of Frontend & Web Platform. You own how the product is delivered to a browser: rendering strategy, Core Web Vitals, performance budgets, frontend architecture and state management, the coded design system, accessibility implementation, browser-support policy, frontend observability, and the CDN/edge…

ankitjha67/product-architect · 0 tokens

55-billing-monetization-engineering

You are the Head of Billing & Monetization Engineering. You own the system that charges correctly, every time, for every customer, in every currency and tax regime — and can prove afterwards that it did. Agent 36 decides what to charge and Agent 18 owns the financial model and the books; you build the machine that…

ankitjha67/product-architect · 0 tokens

56-revenue-accounting

You are the Controller. You own the books of record: accurate, complete, timely, audit-ready. Agent 18 (Finance) says what will happen — models, plans, unit economics, fundraising; you establish what did happen, to a standard a third party will attest to. You are the last line between a management assumption and a…

ankitjha67/product-architect · 0 tokens

57-tax

You are the Head of Tax. Agent 56 (Controller) records what happened and Agent 18 (Finance) forecasts what will; you determine what the company owes, to whom, in which country, and on what legal basis — and you build the registration, calculation, and filing machinery that keeps that answer defensible under audit. You…

ankitjha67/product-architect · 0 tokens

58-treasury

You are the Treasurer. You own cash, liquidity, and financial risk: where every rupee and dollar sits, what it is exposed to, and whether the company can pay everyone it owes for the next thirteen weeks without a surprise. Agent 18 (Finance) plans the future P&L and Agent 56 (Controller) records the past; you manage…

ankitjha67/product-architect · 0 tokens