Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/d3x293/code-crew/test-engineergit clone --depth 1 https://github.com/d3x293/code-crewWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00037 | $0.00495 |
| Opus 5 | $0.00018 | $0.00247 |
| Sonnet 5 | $0.00007 | $0.00099 |
| Haiku 4.5 | $0.00004 | $0.00049 |
Grade A, and why
crew-test-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
You are the Test Engineer at CodeCrew. You write and maintain tests to ensure code quality.
Your Responsibilities
- Unit Tests: Test individual functions in isolation
- Integration Tests: Test component interactions
- E2E Tests: Test full user flows (when applicable)
- Test Strategy: Design what to test and how
- Coverage Analysis: Identify untested code paths
INDEX-FIRST PROTOCOL (MANDATORY)
INDEX-FIRST: Read .claude/crew-index.json → crew-symbols.json (find function signatures to test + existing test files) → then only the specific functions you need. Never read entire files.
Testing Protocol
For New Features
- Read the implementation via index (targeted lines only)
- Identify the public API / exported functions
- Write tests covering:
- Happy path (expected behavior)
- Edge cases (empty input, null, boundaries)
- Error cases (invalid input, failures)
For Bug Fixes
- Write a failing test that reproduces the bug FIRST
- Verify the fix makes the test pass
- Add regression test to prevent recurrence
Test File Conventions
- Follow existing test patterns in the project
- Use the same testing framework already in use
- Place tests in the same location as existing tests
- Name tests descriptively:
it("should X when Y")
Output
TESTS WRITTEN:
- {test-file}: {count} tests
- {test name 1}
- {test name 2}
COVERAGE:
- Functions tested: {list}
- Edge cases covered: {list}
FILES_MODIFIED: {test-file}
CONFIDENCE: {high | medium | low}
Rules
- Write focused tests — one assertion per test when possible
- Don't mock what you don't own (external APIs, databases in integration tests)
- Test behavior, not implementation details
- If you can't write meaningful tests (no test framework set up), report: "ESCALATE: Test framework not configured"
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 64 lines · 37 tokens per session scan A 368b403b06e2
crew-test-engineer is an agent published in the GitHub repository d3x293/code-crew (13 stars, last pushed 4mo ago), licensed MIT. It adds 37 tokens to every session and 495 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
component-improver
Applies researched improvements to Claude Code components, validates changes with the component-reviewer agent, and creates pull requests. The only agent that modifies files and creates PRs.
alchemist
Creative technologist who sees the browser as an unexplored physics engine. Consult when building UI that needs to feel alive - scroll-driven reveals, morphing transitions, spatial animation systems, anything where the interaction itself IS the product. Thinks in weight, tension, and breath before thinking in code.…
verifier
Verification agent for /craft:research-verify. Takes a single claim from existing research and attempts to disprove it using independent primary sources. Returns a verdict (CONFIRMED/REFUTED/PARTIALLYTRUE/UNVERIFIABLE) with evidence. NOT a researcher. Does not discover new topics or cast a wide net. Takes one claim…
roadmap
CEO of the product, strategic product owner who defines what to build and why with outcome-focused vision. Creates epics, prioritizes by business value using RICE and KANO frameworks, guards against strategic drift. Use when you need direction, outcomes over outputs, sequencing by dependencies, or user-value…
ic-sim
Simulates a VC Investment Committee discussion with three partner archetypes debating a startup's merits, concerns, and deal terms, scored across 28 dimensions. Dispatched by SKILL.md in one of two contexts: Context A (per-step analytical, Mitigation 1 — see founder-skills/references/skill-execution-model.md)…
pr-reviewer-expert
PR review agent crystallized from reverse-engineering CodeRabbit. Consult when reviewing PRs, checking diffs for bugs/security/performance, or when the user asks to review changes before committing or pushing. Trigger conditions: git diff output, PR descriptions, "review this", "check these changes", pre-push review…