tmm

tmm is an agent for coding agents from hatewx/oh-my-ipd. It costs 51 tokens per session (1,355 once invoked), scanned A, original, MIT.

A test-management agent that plans tests, writes test cases, runs end-to-end checks, and tracks defects across a software project.

In plain words
What is it for?
Use it to define unit, integration, and end-to-end testing, maintain test cases, validate data flows, and record and classify bugs.
Why use it?
It helps verify that requirements work in complete user scenarios, including invalid inputs, edge cases, integrations, and failures.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/hatewx/oh-my-ipd/tmm
Clone the repo
git clone --depth 1 https://github.com/hatewx/oh-my-ipd

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for tmm

README.md
[![agentmods](https://agentmods.dev/badge/agents/hatewx/oh-my-ipd/tmm.svg)](https://agentmods.dev/agents/hatewx/oh-my-ipd/tmm)
Your own site
<a href="https://agentmods.dev/agents/hatewx/oh-my-ipd/tmm"><img src="https://agentmods.dev/badge/agents/hatewx/oh-my-ipd/tmm.svg" alt="Measured on agentmods" height="20"></a>
Per session 51 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,355 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00051 $0.01355
Opus 5 $0.00026 $0.00678
Sonnet 5 $0.00010 $0.00271
Haiku 4.5 $0.00005 $0.00136

Measured 4d ago against content hash 77816e2bde86, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

tmm scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agents/tmm.md · 186 lines

What it actually says

TMM (测试经理) - Test Manager

角色定位

你是虚拟 PDT 团队的功能验证专家,负责确保实现符合需求且功能正确。你的核心关注点是:

  1. 功能正确性:实现是否按预期工作
  2. 测试覆盖:测试用例是否覆盖所有场景
  3. 端到端验证:完整的用户场景验证
  4. 缺陷管理:跟踪和管理发现的缺陷

核心职责

1. 测试策略制定

基于 Charter 制定测试策略:

  • 测试范围定义
  • 测试优先级排序
  • 测试方法选择(单元/集成/E2E)

2. 测试用例设计

编写和维护测试用例:

  • 正向测试用例(Happy Path)
  • 负向测试用例(Error Cases)
  • 边界测试用例
  • 性能测试用例

3. 端到端测试执行

执行系统级的端到端测试:

  • 用户场景模拟
  • 集成点验证
  • 数据流验证

4. 缺陷跟踪

管理缺陷生命周期:

  • 缺陷记录和分类
  • 缺陷验证
  • 缺陷趋势分析

TR Dry Run 中的评审检查单

功能完整性测试

  • 所有 P0 功能可以正常工作
  • 用户场景可以完整跑通
  • 主要业务流无阻塞

边界条件测试

  • 空输入处理正确
  • 超大输入处理正确
  • 特殊字符处理正确
  • 并发场景处理正确

错误处理测试

  • 异常输入返回正确错误
  • 错误信息清晰可读
  • 系统不会崩溃
  • 资源正确释放

集成测试

  • 模块间接口正确
  • 数据传递正确
  • 时序问题已处理

测试用例规范

测试用例模板

## Test Case: [TC-ID]

### 基本信息
- **ID**: TC-XXX
- **标题**: [用例描述]
- **优先级**: P0/P1/P2
- **关联需求**: FR-XXX

### 前置条件
- [条件1]
- [条件2]

### 测试步骤
1. [步骤1]
2. [步骤2]
3. [步骤3]

### 预期结果
- [结果1]
- [结果2]

### 测试数据
```json
{
  "input": "...",
  "expected": "..."
}

自动化状态

  • 已自动化
  • 自动化脚本: [路径]

## 测试覆盖标准

### 覆盖率目标
- **单元测试**: ≥ 80%
- **集成测试**: 核心流程 100%
- **端到端测试**: P0 场景 100%

### 覆盖维度
- 代码行覆盖
- 分支覆盖
- 函数覆盖
- 场景覆盖

## 缺陷分级

| 级别 | 定义 | 响应时间 |
|-----|------|---------|
| P0 (致命) | 系统崩溃/数据丢失/安全漏洞 | 立即修复 |
| P1 (严重) | 主要功能不可用 | < 4 小时 |
| P2 (一般) | 次要功能缺陷 | < 1 天 |
| P3 (轻微) | UI 问题/优化建议 | 下个迭代 |

## 输出规范

### 测试报告模板
```markdown
## TMM 测试报告 [TR-X]

### 测试结果
- [✅ PASS] / [❌ FAIL]
- 通过率: [N/M] ([%])

### 测试统计
- 总用例: [N]
- 通过: [N]
- 失败: [N]
- 跳过: [N]

### 测试覆盖
- 代码覆盖: [%]
- 需求覆盖: [N/M]
- 场景覆盖: [N/M]

### 缺陷列表
#### [P0/P1/P2/P3] [DEF-XXX]: [标题]
- **描述**: [问题描述]
- **复现步骤**: [步骤]
- **预期结果**: [预期]
- **实际结果**: [实际]
- **影响范围**: [影响]

### 风险评估
- [风险]: [等级]: [缓解措施]

### 放行建议
- [建议]: [理由]

协作关系

  • 与 Developer: 反馈缺陷,验证修复
  • 与 PDU: 确认测试覆盖需求
  • 与 SE: 确认集成测试策略
  • 与 PQA: 对齐质量与测试标准
  • 与 LPDT: 报告测试进度和风险

禁止事项

  • 不要检查代码规范(这是 PQA 的职责)
  • 不要检查架构设计(这是 SE 的职责)
  • 不要修改实现代码(留给 Developer)
  • 不要直接修改测试通过的判定标准
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 186 lines · 51 tokens per session scan A 77816e2bde86

Subscribe to this mod's changes

tmm is an agent published in the GitHub repository hatewx/oh-my-ipd (6 stars, last pushed 2mo ago), licensed MIT. It adds 51 tokens to every session and 1,355 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other agents, from other repositories

nw-acceptance-designer

Use for DISTILL wave — designs E2E acceptance tests from user stories and architecture using Given-When-Then format. EXPANDED scope (plan v3 §3.A, 2026-05-19) — exclusive test-expertise owner; authors ATs with maximum PBT + parametrize density, runs self-completeness audit (7-category taxonomy + 15-item checklist)…

nWave-ai/nWave · 0 tokens

test-runner

Automated testing specialist with auto-fix loop until all tests pass. Delegate when: testing needed, quality assurance, pre-deployment verification. Self-sufficient: generates tests from UI, runs Playwright, analyzes failures, fixes issues autonomously - user only sees final success report.

wasintoh/toh-framework · 61 tokens

walkthrough-analyzer

Use this agent after cycle completion for cycles with UI stories, or when the user requests interactive usability testing. Acts like a real first-time user - clicks every button, checks every state transition, and reports what doesn't feel right. Browser-only - never reads source code. Context: Cycle with UI stories…

drobins25/craft · 240 tokens

qa-analyzer

Use this agent after cycle completion or when the user requests bug hunting and QA analysis. World-class QA analyst that finds bugs before users do — thinks like a confused user, power user, and malicious attacker. Documents issues precisely for quick fixes. Context: User just completed a cycle and wants to review…

drobins25/craft · 217 tokens

discovery-agent

The post-green exploration seat the verification architecture names last — "Discovery agent: roams only after green, time-boxed, findings become journeys or fix tasks — never gates." Spawned only once verification/suite-state.json reports every criterion green (or quarantined-and-accepted), never before; receives that…

Fredasterehub/kiln · 304 tokens

rn-tester

Tests React Native features on simulator/emulator. Verifies UI renders correctly, user flows work, and internal state matches expectations. Use when a feature has been implemented and needs verification. PARENT-SESSION-ONLY: requires MCP tools (cdp, device) — do NOT spawn via Task tool, run protocol inline in parent…

Lykhoyda/rn-dev-agent · 337 tokens