Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add williamzujkowski/standards --skill testinggit clone --depth 1 https://github.com/williamzujkowski/standardsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/williamzujkowski/standards/testing)<a href="https://agentmods.dev/skills/williamzujkowski/standards/testing"><img src="https://agentmods.dev/badge/skills/williamzujkowski/standards/testing/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/williamzujkowski/standards/testing"><img src="https://agentmods.dev/badge/skills/williamzujkowski/standards/testing.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00021 | $0.04267 |
| Opus 5.5 | $0.00008 | $0.01707 |
| Sonnet 5.5 | $0.00004 | $0.00853 |
| Haiku 4.5 | $0.00002 | $0.00427 |
Grade A, and why
testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 736 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Testing Standards Skill
Level 1: Quick Start (5 minutes)
What You'll Learn
Implement effective testing strategies to ensure code quality, catch bugs early, and maintain confidence in changes.
Core Principles
- Test First: Write tests before implementation (TDD)
- Coverage: Aim for >80% code coverage
- Isolation: Tests should be independent and repeatable
- Speed: Keep unit tests fast (<100ms each)
Testing Pyramid
/\
/ \ E2E Tests (Few)
/----\
/ Inte \
/ gration \ (Some)
/ Tests \
/------------\
/ Unit Tests \ (Many)
/--------------\
Quick Start Example
// 1. Write the test first (RED)
describe("calculateDiscount", () => {
it("should apply 10% discount for premium users", () => {
const user = { tier: "premium" };
const order = { total: 100 };
const result = calculateDiscount(user, order);
expect(result).toBe(90);
});
});
// 2. Implement minimal code (GREEN)
function calculateDiscount(user, order) {
if (user.tier === "premium") {
return order.total * 0.9;
}
return order.total;
}
// 3. Refactor (REFACTOR)
function calculateDiscount(
user: User,
order: Order
): number {
const PREMIUM_DISCOUNT = 0.1;
return user.tier === "premium"
? order.total * (1 - PREMIUM_DISCOUNT)
: order.total;
}
Essential Checklist
- Unit tests for all business logic
- Integration tests for critical paths
- Tests run in CI/CD pipeline
- Coverage >80%
- Tests are maintainable and readable
Common Pitfalls
- Testing implementation details instead of behavior
- Slow tests that discourage running them
- Flaky tests that fail intermittently
- Poor test isolation (shared state)
Level 2: Implementation (30 minutes)
Deep Dive Topics
1. Test-Driven Development (TDD)
Red-Green-Refactor Cycle:
// STEP 1: RED - Write failing test
describe("UserService", () => {
describe("createUser", () => {
it("should hash password before saving", async () => {
const userData = {
email: "[email protected]",
password: "SecurePass123!",
};
const service = new UserService(mockRepository, mockHasher);
await service.createUser(userData);
expect(mockHasher.hash).toHaveBeenCalledWith("SecurePass123!");
expect(mockRepository.save).toHaveBeenCalledWith(
expect.objectContaining({
email: "[email protected]",
passwordHash: expect.any(String),
})
);
});
});
});
// STEP 2: GREEN - Minimal implementation
class UserService {
constructor(
private repository: UserRepository,
private hasher: PasswordHasher
) {}
async createUser(userData: UserData): Promise<User> {
const passwordHash = await this.hasher.hash(userData.password);
return this.repository.save({
...userData,
passwordHash,
});
}
}
// STEP 3: REFACTOR - Improve design
class UserService {
constructor(
private repository: UserRepository,
private hasher: PasswordHasher,
private validator: UserValidator
) {}
async createUser(userData: UserData): Promise<User> {
await this.validator.validate(userData);
const passwordHash = await this.hasher.hash(userData.password);
const user = new User({
email: userData.email,
passwordHash,
createdAt: new Date(),
});
return this.repository.save(user);
}
}
What ships with it
48 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- e2e-testing/config/cypress.config.ts 3.7 KB runs code
- e2e-testing/config/playwright.config.ts 4.1 KB runs code
- e2e-testing/config/README.md 928 B
- e2e-testing/REFERENCE.md 39 KB
- e2e-testing/resources/e2e-best-practices.md 15 KB
- e2e-testing/resources/README.md 962 B
- e2e-testing/scripts/README.md 1.2 KB
- e2e-testing/scripts/run-e2e-tests.sh 7.0 KB runs code
- e2e-testing/SKILL.md 10 KB
- e2e-testing/templates/page-object.ts 11 KB runs code
- e2e-testing/templates/README.md 943 B
- e2e-testing/templates/test-template.spec.ts 11 KB runs code
- integration-testing/resources/examples/contract-testing.md 6.4 KB
- integration-testing/resources/README.md 36 B
- integration-testing/scripts/README.md 34 B
- integration-testing/scripts/setup-test-db.sh 1.9 KB runs code
- integration-testing/SKILL.md 14 KB
- integration-testing/templates/api-test-template.js 3.3 KB runs code
- integration-testing/templates/docker-compose.test.yml 1.1 KB
- integration-testing/templates/integration-test-python.py 4.2 KB runs code
- integration-testing/templates/README.md 36 B
- performance-testing/config/jmeter-test-plan.jmx 8.1 KB
- performance-testing/REFERENCE.md 16 KB
- performance-testing/resources/performance-checklist.md 8.1 KB
- performance-testing/resources/README.md 36 B
- performance-testing/scripts/README.md 34 B
- performance-testing/scripts/run-perf-tests.sh 3.8 KB runs code
- performance-testing/SKILL.md 10 KB
- performance-testing/templates/grafana-dashboard.json 2.3 KB
- performance-testing/templates/k6-load-test.js 9.9 KB runs code
- performance-testing/templates/k6-stress-test.js 2.0 KB runs code
- performance-testing/templates/README.md 36 B
- unit-testing/config/jest.config.js 7.0 KB runs code
- unit-testing/config/pytest.ini 6.3 KB
- unit-testing/resources/configs/jest.config.js 1.2 KB runs code
- unit-testing/resources/configs/pytest.ini 1.1 KB
- unit-testing/resources/coverage-configs/.coveragerc 587 B
- unit-testing/resources/README.md 29 B
- unit-testing/resources/test-pyramid-diagram.md 7.0 KB
- unit-testing/scripts/README.md 27 B
- unit-testing/SKILL.md 17 KB
- unit-testing/templates/example_test.go 11 KB
- unit-testing/templates/example.test.js 11 KB runs code
- unit-testing/templates/README.md 29 B
- unit-testing/templates/test_example.py 9.8 KB runs code
- unit-testing/templates/test-template-go.go 1.6 KB
- unit-testing/templates/test-template-jest.js 2.9 KB runs code
- unit-testing/templates/test-template-pytest.py 2.9 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 736 lines · 21 tokens per session scan A 1959fc571e54
testing is a skill published in the GitHub repository williamzujkowski/standards (18 stars, last pushed 1mo ago), licensed MIT. It adds 21 tokens to every session and 4,267 once invoked, about $0.0001 per session on Opus 5.5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-29.
Other skills, from other repositories
engram-testing-coverage
TDD and coverage standards for Engram. Trigger: When implementing behavior changes in any package.
iterative-development
TDD iteration loops using Claude Code Stop hooks - runs tests after each response, feeds failures back automatically.
python
Python development with ruff, mypy, pytest - TDD and type safety.
mobiai-mobile-tdd
You MUST use this before writing any implementation code for a mobile feature, bug fix, refactor, or behavior change. Tests come before implementation — no exceptions.
tdd
This skill should be used when the user wants to implement features or fix bugs using test-driven development. Enforces the RED-GREEN-REFACTOR cycle with vertical slicing, context isolation between test writing and implementation, human checkpoints, and auto-test feedback loops. Uses multi-agent orchestration with the…
nw-fp-clojure
Clojure language-specific patterns, data-first modeling, REPL-driven development, and spec.