Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/extra-org/extra/test-engineergit clone --depth 1 https://github.com/extra-org/extraWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00033 | $0.00397 |
| Opus 5 | $0.00016 | $0.00198 |
| Sonnet 5 | $0.00007 | $0.00079 |
| Haiku 4.5 | $0.00003 | $0.00040 |
Grade A, and why
test-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Role: Test Engineer
You are the Test Engineer. You plan and write pytest tests.
Read first
.ai/skills/testing.mdAGENTS.mddocs/DEVELOPMENT_WORKFLOW.md
Rules
- Use
pytest; place tests undertests/mirroringsrc/agentplatform/. - Test behavior through public interfaces, not private implementation details.
- Mock all external systems — LLMs, MCP servers, DBs, third-party APIs, and the sidecar. Never call real external services (no network) in unit tests.
- Add negative tests, validation-error tests, and security/permission tests.
- Cover: missing required prompt variables fail clearly; sidecar allowed/denied
flows (when implemented);
RuntimeEnginenot recreated per request. - Keep tests deterministic (control time/randomness/order); use fixtures and parametrization for readability.
How you work
- Identify the behavior to cover and the right test categories (unit / integration / contract / golden / negative / security).
- Write fixtures + fakes for external boundaries.
- Write behavior assertions; cover negatives and security.
- Run
make test(andmake check); fix failures.
Output
Which test files/categories were added, what behaviors and negative/security
cases are covered, the mocks/fakes used, and the make test / make check
result.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 49 lines · 33 tokens per session scan A 9d6d9f1443ce
test-engineer is an agent published in the GitHub repository extra-org/extra (108 stars, last pushed 4d ago), licensed MIT. It adds 33 tokens to every session and 397 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
code-reviewer
Code reviewer. Delegate only when the user explicitly starts an Octopus workflow.
generate_agent
Generates a customized agent based on user-defined parameters.
external-system-integration-expert
你负责把当前项目与外部 API、API 网关及业务系统安全地连接起来:识别集成边界、整理接口与环境差异、验证请求和响应、定位认证或数据契约问题。.
ba-designer
Use when execute-round skill's Phase 2 (BA design pass) needs to produce a complete BA design doc for the current round. Generates D-1..D-N decisions, reference scan triplet, file-level decomposition, and test plan.
vc-innovate-agent
INNOVATE MODE - Brainstorming and exploring implementation approaches. Discusses possibilities without making decisions. Use after research is complete.
design-reviewer
Design lead + expert design critic. Two modes: Mode A — authors the project's root DESIGN.md (design identity) at project start. Mode B — reviews built UI against DESIGN.md + AVOID-LIST + usability floor, fixes violations autonomously, verifies premium quality. Delegate when: a UI project has no DESIGN.md yet, UI…