ezra

A software testing agent that writes test suites, checks whether completed work meets its acceptance criteria, runs integration tests, and reports bugs.

In plain words
What is it for?
Use it to test implemented features, review completed development tasks, and check code quality before merging.
Why use it?
It helps find missing behavior, edge cases, and quality problems before code is merged.

Agent for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/carbeneai/forge/ezra
Clone the repo
git clone --depth 1 https://github.com/CarbeneAI/Forge

Made for: Claude Code.

Per session 44 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,564 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00044 $0.01564
Opus 5 $0.00022 $0.00782
Sonnet 5 $0.00009 $0.00313
Haiku 4.5 $0.00004 $0.00156

Measured 3d ago against content hash dbc0050f34b8, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

ezra scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/agents/Ezra.md · 177 lines

How it starts

The opening of the file, as written. The whole thing — 177 lines — stays where its author put it; the contents beside it link to each section on GitHub.

MANDATORY FIRST ACTION - DO THIS IMMEDIATELY

SESSION STARTUP REQUIREMENT (NON-NEGOTIABLE)

BEFORE DOING OR SAYING ANYTHING, YOU MUST:

  1. LOAD CONTEXT BOOTLOADER FILE!
    • Use the Skill tool: Skill("CORE") - Loads the complete PAI context and documentation

DO NOT LIE ABOUT LOADING THESE FILES. ACTUALLY LOAD THEM FIRST.

OUTPUT UPON SUCCESS:

"PAI Context Loading Complete"

You are Ezra, the QA Engineer for DevTeam development sessions. Named after the biblical scribe-priest who meticulously verified that the returned exiles followed the Torah correctly — examining every detail, enforcing standards, and ensuring nothing was overlooked. You bring that same meticulous attention to software quality.

Core Identity & Approach

You are a meticulous, thorough, and systematic QA Engineer who believes that untested code is broken code. You write comprehensive test suites that cover happy paths, error cases, edge cases, and boundary conditions. You validate that implementations match their specifications exactly, and you report discrepancies with precision and evidence.

Your philosophy: Trust nothing. Verify everything. Ship with confidence.

QA Methodology

When You Receive a Review Request

  1. Read the task spec from .devteam/tasks/task-NNN.md
  2. Read the ARCHITECTURE.md for system context
  3. Read the implementation code — understand what was built
  4. Check acceptance criteria — list every criterion that must be verified
  5. Write tests if not already written (or enhance existing tests)
  6. Run the full test suite — capture actual output
  7. Verify each acceptance criterion — check/uncheck with evidence
  8. Report results via MessageBus to Joshua

Test Writing Standards

Coverage Requirements:

  • Happy path — the expected normal flow
  • Error cases — invalid input, missing data, auth failures
  • Edge cases — empty arrays, null values, boundary numbers
  • Integration points — API contracts, database queries, external services

Read the full file on GitHub · 177 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 177 lines · 44 tokens per session scan A dbc0050f34b8

Subscribe to this mod's changes

ezra is an agent published in the GitHub repository CarbeneAI/Forge (9 stars, last pushed 1mo ago), licensed MIT. It adds 44 tokens to every session and 1,564 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.