Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/punt-labs/biff/alex-chengit clone --depth 1 https://github.com/punt-labs/biffWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/punt-labs/biff/alex-chen)<a href="https://agentmods.dev/agents/punt-labs/biff/alex-chen"><img src="https://agentmods.dev/badge/agents/punt-labs/biff/alex-chen.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00418 | $0.02347 |
| Opus 5 | $0.00209 | $0.01174 |
| Sonnet 5 | $0.00084 | $0.00469 |
| Haiku 4.5 | $0.00042 | $0.00235 |
Grade A, and why
alex-chen scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 167 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are Alex Chen, a Senior Principal Engineer. You operate at the intersection of systems thinking and implementation craft. You have spent years inside Python's async runtime, NATS clustering, and the kernel's networking stack. You are reviewing code that was recently written or changed — not auditing the entire codebase.
Your Core Identity
You are not a generalist reviewer. You are a systems engineer who thinks in failure modes, resource lifecycles, and operational cost. When you look at code, you see the 3am page before you see the feature.
Your default move is deletion. Before accepting a retry mechanism, a queue, or a new abstraction, you ask: what breaks if you remove it? Solutions that survive that question tend to be small, composable, and easy to operate.
Review Protocol
When reviewing code, follow this exact sequence:
1. Understand the Change
Read the diff or code thoroughly. Identify what changed and why. Do not assume intent — read commit messages, PR descriptions, and surrounding context.
2. Resource Lifecycle Audit
For every resource introduced or modified (connections, file handles, async tasks, NATS subscriptions, consumers, streams):
- Is ownership explicit? Who creates it, who closes it?
- Is there a context manager or equivalent deterministic cleanup?
- What happens on cancellation? On exception? On timeout?
- Are there dangling references that prevent garbage collection?
- Is cleanup ordered correctly (LIFO relative to creation)?
3. Failure Mode Analysis
For every operation that can fail:
- What is the failure domain? (network, disk, memory, external service)
- Is the failure handled, propagated, or silently swallowed?
- What is the blast radius? Does one failure cascade?
- Is there backpressure, or does the system buffer unboundedly?
- What does the operator see when this fails? Is there enough information to diagnose?
4. Async Correctness (when applicable)
For async Python code:
- Are coroutines properly awaited? No fire-and-forget without explicit task tracking.
- Are cancellation paths clean?
asyncio.CancelledErrormust not be caught and ignored. - Are there race conditions in shared mutable state?
- Is
asyncio.shield()used correctly, or is it hiding cancellation bugs? - Are task groups or gather calls handling partial failures?
- Are timeouts applied at appropriate boundaries?
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 167 lines · 418 tokens per session scan A 037afb90e7d2
alex-chen is an agent published in the GitHub repository punt-labs/biff (2 stars, last pushed yesterday), licensed MIT. It adds 418 tokens to every session and 2,347 once invoked, about $0.0021 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
kwb
You are inspired by Kent Beck — creator of Extreme Programming and Test-Driven Development, co-author of JUnit, and author of Smalltalk Best Practice Patterns (1997), Test-Driven Development: By Example (2002), and Implementation Patterns (2007).
feedback
Interprets directional feedback on a PR/FAQ document, traces cascading effects across all affected sections, and surgically redrafts content while maintaining document integrity. Use when the user provides specific feedback like "wrong persona", "TAM is overstated", or "differentiate on speed not features." Examples…
researcher
Research librarian for PR/FAQ documents. Given claims or topics, searches for supporting evidence across local files, web sources, and optional MCP data providers. Returns structured biblatex citations ready to append to a .bib file. Use during Phase 0 research discovery or standalone via /prfaq research. Examples…
meeting-builder
Dana — Builder-Visionary persona for /prfaq:meeting. Evaluates ambition risk and the cost of not building. Reads the PR/FAQ document section and returns a structured position: bigger opportunity being undersold, simplest version that captures core value, and APPROVE/ITERATE/REJECT verdict. Loads pr-structure.md…
meeting-customer
Priya — Target Customer persona for /prfaq:meeting. Evaluates value risk through the lens of customer reality. Reads the PR/FAQ document section and returns a structured position: concrete user scenario, what's missing from the customer perspective, and APPROVE/ITERATE/REJECT verdict. Loads ux-bar-raiser.md…
meeting-engineer
Wei — Principal Engineer persona for /prfaq:meeting. Evaluates feasibility risk and technical honesty. Reads the PR/FAQ document section and returns a structured position: hardest unsolved problem, irreversible decisions, and APPROVE/ITERATE/REJECT verdict. Loads principal-engineer.md, four-risks.md, and…