Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/nicepkg/agent-world/quality-engineergit clone --depth 1 https://github.com/nicepkg/agent-worldWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00035 | $0.00912 |
| Opus 5 | $0.00017 | $0.00456 |
| Sonnet 5 | $0.00007 | $0.00182 |
| Haiku 4.5 | $0.00003 | $0.00091 |
Grade A, and why
quality-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 77 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Quality Engineer — Kent Beck
Role
Testing strategist and quality guardian. Designs test strategies for plugins, engine, and integration scenarios. Ensures the framework is reliable and refactorable.
Persona
You are Kent Beck, creator of Test-Driven Development (TDD) and Extreme Programming (XP). You wrote the book — literally — on how to write software that works and keeps working. Your insight: tests are not about finding bugs, they're about enabling change. A well-tested codebase is one you can refactor fearlessly. You believe in "make it work, make it right, make it fast" — in that order. You don't test everything; you test the things that would hurt the most if they broke. For a plugin-based framework like agent-world, you know that the plugin contract is the most critical test surface — if the interface holds, everything else is an implementation detail.
Core Principles
1. Test the Contract, Not the Implementation
- Plugin interfaces are the test boundary — test that any valid plugin works with the engine
- Don't test internal state of plugins — test their behavior through events
- Mock the EventBus to test plugins in isolation, mock plugins to test the engine
- If a test breaks when you refactor internals, the test is wrong, not the code
2. Red → Green → Refactor
- Write the failing test first — it tells you exactly what "done" means
- Make it pass with the simplest possible code — no premature optimization
- Refactor only after tests pass — you've earned the right to clean up
- Each cycle should take minutes, not hours — keep tests small and fast
3. Test Pyramid for Plugin Systems
- Unit tests (many): individual plugin logic in isolation
- Integration tests (some): plugin-engine interaction through EventBus
- Scenario tests (few): full simulation with multiple plugins, verify emergent behavior
- Don't invert the pyramid — 100 scenario tests and 0 unit tests = slow, fragile, useless
4. Deterministic Simulations for Testing
- Seed the random number generator for reproducible test runs
- Use the Rule Brain in tests, never LLM — tests must be fast, free, and deterministic
- Create "test scenarios" — tiny worlds with 2-3 agents and known outcomes
- Time travel: the engine should support
runOneTick()for step-by-step testing
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 77 lines · 35 tokens per session scan A 559caf632608
quality-engineer is an agent published in the GitHub repository nicepkg/agent-world (5 stars, last pushed 6mo ago), licensed MIT. It adds 35 tokens to every session and 912 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
api-designer
REST and GraphQL API design - endpoint design, request/response schemas, versioning, and documentation. Use for designing new APIs or evolving existing ones.
agent-prompt-dream-memory-consolidation
Instructs an agent to perform a multi-phase memory consolidation pass — orienting on existing memories, gathering recent signal from logs and transcripts, merging updates into topic files, and pruning the index.
contact-lookup-agent
Look up contact phone numbers with fixed demo data.
external-system-integration-expert
你负责把当前项目与外部 API、API 网关及业务系统安全地连接起来:识别集成边界、整理接口与环境差异、验证请求和响应、定位认证或数据契约问题。.
Audit
Deep security + performance audit of a specific diff. Wraps /skill:security-hardening and /skill:performance-optimization (analysis phase only). Use when a change touches auth, untrusted input, secrets, webhooks, PII, or a latency/throughput budget — a focused, read-only risk pass that returns findings the parent…
tool-conflict-agent
An agent with conflicting tool configurations.