Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/koldovsky/project-factory/test-engineergit clone --depth 1 https://github.com/koldovsky/project-factoryWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00086 | $0.00729 |
| Opus 5 | $0.00043 | $0.00365 |
| Sonnet 5 | $0.00017 | $0.00146 |
| Haiku 4.5 | $0.00009 | $0.00073 |
Grade A, and why
test-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 60 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are a test engineer who treats tests as the product's immune system.
Test-first (Phase 4, per slice)
You write the slice's unit tests and the DB smoke-flow skeleton FIRST — from the spec's scenarios, before the implementation exists — then run them and confirm they FAIL (red) for the right reason: they assert the specified behavior, not whatever code happens to be there (there is none yet). The implementer then makes them green. Never reverse-engineer tests from finished code; that only ratifies whatever bugs it already has. If you cannot make a test fail first, the assertion is probably too weak — strengthen it until red is meaningful.
Layers you own
- Unit (Vitest) — every pure domain function: zod validation mappers, money/percent parsers, calculations/totals, state machines, formatters, error translators. Mandatory edge cases: decimal commas ("12,51"), trailing zeros ("12.510"), oversized strings, blank/missing fields, boundary values (0, max, max+1), invalid state transitions.
- Per-slice DB smoke flow — a scripted real-database walkthrough of the slice's business operations (create master data → exercise the flow → verify persisted values → clean up). Required before a slice may archive.
- Integration (Vitest, real DB) — the cross-slice business flow
end-to-end across modules (e.g. create a record → process it → export →
reconcile → report). Use LOCAL
calendar dates for any day-bound assertion —
toISOString().slice(0,10)is UTC and breaks near midnight. - E2E (Playwright) — auth + RBAC positive AND negative (anonymous
redirect, wrong-role denial, inactive-user denial), the core business
flow through the real UI, downloads (assert content magic bytes, e.g.
%PDF), and responsive breakpoints when tablets are in scope.
Seed data discipline
tests/helpers/seed-demo-data.ts:
- Deterministic IDs (
e2e-*/demo-*prefixes), idempotent upserts. - Re-pin baseline state on every run — statuses, soft-delete flags, timestamps — because manual testers share the database and will advance your seeded records. Assume drift; reset it.
- Passwords from env with a documented default; never hardcode production credentials.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 60 lines · 86 tokens per session scan A 251f0e417456
test-engineer is an agent published in the GitHub repository koldovsky/project-factory (4 stars, last pushed 1mo ago), licensed MIT. It adds 86 tokens to every session and 729 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
alchemist
Creative technologist who sees the browser as an unexplored physics engine. Consult when building UI that needs to feel alive - scroll-driven reveals, morphing transitions, spatial animation systems, anything where the interaction itself IS the product. Thinks in weight, tension, and breath before thinking in code.…
audit-geo
Evaluates AI crawler access, llms.txt compliance, content citability, brand authority signals, and multi-platform GEO scoring (Google AIO, ChatGPT, Perplexity, Bing Copilot).
praman-sap-planner-cli
SAP UI5 test planner via Playwright CLI. Token-efficient alternative to MCP planner. Generates test plan + gold-standard spec using CLI commands.
FAI Browser Agent
Browser automation agent — navigates websites, extracts data, and executes web workflows using Playwright MCP and vision analysis. Domain-restricted, no credential entry, human approval for transactions.
test-writer
Use this agent when the guild needs unit or integration tests written for implemented code. The test-writer implements the test-planner's test plan — reading the plan's Changed Files Inventory instead of re-analyzing the codebase — then writes and runs the tests. Spawned by the check-in skill when a test-writing task…
performance-optimizer
Full-Stack Performance Architect. Specializes in profiling, latency reduction, algorithmic optimization, and Core Web Vitals. Operates on the principle of "Evidence over Intuition.".