Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/johanthoren/jeff/testingnpx skills add johanthoren/jeff --skill testinggit clone --depth 1 https://github.com/johanthoren/jeffWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00079 | $0.01258 |
| Opus 5 | $0.00039 | $0.00629 |
| Sonnet 5 | $0.00016 | $0.00252 |
| Haiku 4.5 | $0.00008 | $0.00126 |
Grade A, and why
testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 62 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Testing Standards
Language-agnostic defaults. Language-specific skills override where they conflict.
Golden Rule: If you can't test it easily, refactor it.
Structure and Naming
- Arrange, act, assert: set up the data, execute the code, verify the result. One behavior per test.
- Names state the expectation:
validateEmail returns false for invalid format, notit worksortest user.
What to Test
- DO: happy path; edge cases (boundaries, empty, null); error cases (invalid input, failures); business logic; public APIs.
- DON'T: third-party libraries; framework internals; simple getters/setters; private implementation details.
Coverage
Coverage is an outcome of testing the right intersections, not a target to chase. Not every line, and not every one-liner, needs a test; a one-line pass-through does not. Cover each behavior and its real edges; a task's acceptance criteria are the floor.
Principles
- Test behavior, not implementation: focus on what, not how. Assert results, never that a mock was called with particular arguments:
expect(mock).toHaveBeenCalledWith(...)pins procedure and goes red on a behavior-preserving refactor. - Would it survive a refactor? If a behavior-preserving refactor would turn the test red, it tests procedure; rewrite it to assert the result. Procedure-coupled tests lock in design flaws.
- Mock through injected boundaries: dependency injection makes a hand-rolled fake (
{ findById: () => fixture }) sufficient; no framework magic required. - Independent tests: no shared state, any order, one assertion's worth of behavior each.
- Fast and reliable: run tests frequently; fix failures immediately.
Change-Detector Tests Are a Banned Smell
A change-detector test asserts a value or call shape that no consumer observes, so it goes red on any edit to that value and catches no regression the edit would not. It is configuration duplicated into the test as a second place to edit. Do not write one.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 62 lines · 79 tokens per session scan A cdf8fcfb06a8
testing is a skill published in the GitHub repository johanthoren/jeff (4 stars, last pushed 5d ago), licensed Apache-2.0. It adds 79 tokens to every session and 1,258 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
agent-code-analyzer
Agent skill for code-analyzer - invoke with $agent-code-analyzer.
foundry-config-setup
Resolve missing setup caused by a hardcoded Foundry project endpoint or model in a sample. Use when a sample fails because it uses a placeholder/hardcoded projectendpoint (for example "https://your-project.services.ai.azure.com") or a hardcoded model instead of reading them from the environment.
evolve
Start or monitor an evolutionary development loop.
agile-product-owner
../../../product-team/agile-product-owner/skills/agile-product-owner/SKILL.md.
agent-memory
../../../engineering/agent-memory/skills/agent-memory/SKILL.md.
dogfood
Systematically explore and test a mobile app on iOS/Android with agent-device to find bugs, UX issues, and other problems. Use when asked to dogfood, QA, exploratory test, find issues, bug hunt, or test this app on mobile.