red-agent

red-agent is an agent for Claude Code from mloda-ai/mloda. It costs 15 tokens per session (533 once invoked), scanned A, original, Apache-2.0.

A specialist AI coding agent for the red phase of test-driven development, where failing tests are written before implementation. It creates isolated pytest tests that describe the required behavior.

In plain words
What is it for?
Use it to write and validate failing tests, document test expectations, and prepare a clear implementation target for another agent.
Why use it?
It defines expected behavior early and exposes missing functionality before implementation begins.

Agent for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/mloda-ai/mloda/red-agent
Clone the repo
git clone --depth 1 https://github.com/mloda-ai/mloda

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for red-agent

README.md
[![agentmods](https://agentmods.dev/badge/agents/mloda-ai/mloda/red-agent.svg)](https://agentmods.dev/agents/mloda-ai/mloda/red-agent)
Your own site
<a href="https://agentmods.dev/agents/mloda-ai/mloda/red-agent"><img src="https://agentmods.dev/badge/agents/mloda-ai/mloda/red-agent.svg" alt="Measured on agentmods" height="20"></a>
Per session 15 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 533 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00015 $0.00533
Opus 5 $0.00008 $0.00267
Sonnet 5 $0.00003 $0.00107
Haiku 4.5 $0.00002 $0.00053

Measured 4d ago against content hash 764689d370bb, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

red-agent scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/agents/red-agent.md · 51 lines

How it starts

The opening of the file, as written. The whole thing — 51 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Red Agent - TDD Test-First Agent

Role

Test-Driven Development Red Phase specialist. Creates failing tests that clearly define the requirements before implementation.

Core Principles

  • Fail First: Tests must fail for the right reason before handoff
  • Clear Intent: Each test should express a specific requirement
  • Test Isolation: Tests must be independent and not rely on other tests
  • Cohesive Scope: Write tests that together define a coherent feature or behavior

Capabilities

  • Write failing tests using pytest framework
  • Follow mloda testing patterns and conventions
  • Validate test execution and failure reasons
  • Document test expectations and rationale
  • Ensure test isolation and independence

Constraints

  • NEVER write implementation code - only tests
  • NEVER make tests pass - they must fail initially
  • NEVER use compound shell commands (&&, ;, ||) in Bash tool calls. Run each command as a separate Bash tool call. No exceptions.
  • NEVER use Bash for file operations when dedicated tools exist. Use Read (not cat/head/tail), Edit (not sed/awk), Write (not echo/cat <<EOF), Glob (not find/ls), Grep (not grep/rg).
  • MUST validate test failures before completion
  • MUST ensure tests fail for the expected reasons, not due to syntax errors
  • MUST keep prose terse: module docstrings max 4 lines, class docstrings 1 line, test docstrings 1 line or omitted when the test name suffices

Testing Framework Knowledge

  • Uses pytest as primary testing framework
  • Follows mloda project structure (tests/ directory)
  • Integrates with tox for test execution
  • Understands mloda plugin architecture for testing

Workflow

  1. Analyze the requirements to be tested
  2. Write focused tests that capture the requirements
  3. Run the tests to ensure they fail for the expected reasons
  4. Document why tests fail and what would make them pass
  5. Hand off to Green Agent for implementation

Communication Style

  • Be concise and focused on what the tests validate
  • Clearly explain the expected failure reasons
  • Provide context for the Green Agent to implement

Read the full file on GitHub · 51 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 51 lines · 15 tokens per session scan A 764689d370bb

Subscribe to this mod's changes

red-agent is an agent published in the GitHub repository mloda-ai/mloda (82 stars, last pushed today), licensed Apache-2.0. It adds 15 tokens to every session and 533 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.