Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add OKHP3/skillz --skill integration-testinggit clone --depth 1 https://github.com/OKHP3/skillzWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/okhp3/skillz/integration-testing)<a href="https://agentmods.dev/skills/okhp3/skillz/integration-testing"><img src="https://agentmods.dev/badge/skills/okhp3/skillz/integration-testing/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/okhp3/skillz/integration-testing"><img src="https://agentmods.dev/badge/skills/okhp3/skillz/integration-testing.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00078 | $0.00966 |
| Opus 5 | $0.00039 | $0.00483 |
| Sonnet 5 | $0.00016 | $0.00193 |
| Haiku 4.5 | $0.00008 | $0.00097 |
Grade A, and why
integration-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 93 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Integration testing
Unit tests prove each piece works against your assumptions about its neighbours. Integration tests prove the assumptions were right. Almost every "all tests pass but production is broken" story lives in that gap — a mock that returns a shape the real service never returns.
These cost more than unit tests and less than end-to-end. Spend them on the seams where your assumptions about someone else's behaviour are most likely wrong.
1. Pick the seams worth the cost
Test the boundaries where you cross into something you do not control:
- Your code against a real database: queries, migrations, transactions, constraints. The highest-value integration tests in most applications, because SQL and ORM behaviour is where mocks lie most.
- Your service against its dependencies: real HTTP, real serialization.
- Message producers against consumers: that the payload one writes is one the other parses.
- Your code against the filesystem, clock, or queue, where behaviour is subtle.
Do not integration-test pure logic. If it has no boundary, it belongs in a unit test.
Done when: each planned test crosses a boundary you do not own.
2. Use the real thing, not a mock
The entire point is exercising real behaviour. A mocked database tests your mock.
- Containers for infrastructure: a real Postgres, Redis, or Kafka in a container. Testcontainers-style libraries make this a few lines and it is worth it.
- The same version as production. Testing on a different major version tests a different system.
- Real HTTP against a local instance where you can run the dependency; a recorded or contract-verified stub where you cannot.
For third-party services you cannot run: record real responses once, replay them, and re-record on a schedule. A hand-written stub drifts from reality silently and gives false confidence indefinitely.
Done when: no test in this suite mocks the thing it is testing against.
3. Make each test own its data
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 93 lines · 78 tokens per session scan A e1c3c8d5f035
integration-testing is a skill published in the GitHub repository OKHP3/skillz (3 stars, last pushed yesterday), licensed MIT. It adds 78 tokens to every session and 966 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
kelly-agent-eval
Review board (Busabase App-in-Skill) that runs a fixed suite of mock test cases against a baseline vs candidate agent version and surfaces rubric-scored regressions before a release. Use when the user invokes $kelly-agent-eval or /kelly-agent-eval, wants to review agent-version regressions, compare baseline vs…
kelly-app-skill-creator-tests
Build, maintain, and run conformance tests for canonical App-in-Skill projects created by kelly-app-skill-creator. Use when a Kelly app skill needs contract checks, local server smoke tests, responsive browser acceptance, temporary open-source Busabase integration, environment-gated Busabase Cloud OAuth verification…
pentest-api-attacker
Test APIs against OWASP API Security Top 10 including discovery, auth abuse, and protocol-specific checks.
pentest-container-k8s
Test Docker and Kubernetes security controls for RBAC abuse, breakout, and secret exposure.
pentest-remediation-validator
Retest remediated findings, detect regressions, and generate remediation status and certification artifacts.
pentest-vuln-analyzer
Correlate scanner results with CVE and exploit intelligence and prioritize by CVSS and exploitability.