Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/kid-sid/codex-spellbook/integration-testingnpx skills add kid-sid/codex-spellbook --skill integration-testinggit clone --depth 1 https://github.com/kid-sid/codex-spellbookWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00034 | $0.00629 |
| Opus 5 | $0.00017 | $0.00315 |
| Sonnet 5 | $0.00007 | $0.00126 |
| Haiku 4.5 | $0.00003 | $0.00063 |
Grade A, and why
integration-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 78 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Integration Testing
Verify the interaction between components, data stores, and external services to ensure contract correctness and state integrity.
When to Activate
- Test database interactions and repository logic
- Verify API client behavior against wire services
- Ensure middleware and handlers work together
- Test multi-service flows in a single process
- Debug cross-component state synchronization
- Refactor internals while keeping public contracts stable
- Audit data persistence and retrieval behavior
Testing Layers
| Level | Isolation | Speed | Confidence |
|---|---|---|---|
| Unit | Pure functions, mocked collaborators | Fastest | Low (Logic only) |
| Integration | Real database, real network clients | Medium | Medium (Contracts + State) |
| E2E | Real browser, full environment | Slowest | High (User journey) |
Database Integration
| Strategy | When to use | Rule |
|---|---|---|
| Per-test Transaction | Standard relational DB tests | Wrap in begin, then rollback on teardown |
| Unique Fixtures | Parallel tests or non-transactional DBs | Use UUIDs/Random prefixes forทุก record |
| Truncation | Between suites or for dirty state | Truncate only the tables used |
BAD
def test_get_user_orders():
# Relies on order ID 1 existing from previous test
orders = get_user_orders(user_id=1)
assert len(orders) == 1
GOOD
def test_get_user_orders(db_session, user_factory, order_factory):
user = user_factory.create()
order_factory.create(user_id=user.id)
orders = get_user_orders(user_id=user.id)
assert len(orders) == 1
API Boundaries
| Concern | Strategy |
|---|---|
| HTTP Clients | Use WireMock or Prism instead of mocking the client class |
| Message Brokers | Use a real containerized instance (Testcontainers) |
| File Storage | Use a local filesystem fake or MinIO |
Checklist
- Database tests use a dedicated clean environment (e.g. Testcontainers)
- Tests create their own data rather than relying on shared state
- Teardown logic reliably cleans up files, records, and connections
- Network calls to external APIs are intercepted at the wire level
- Tests verify side effects like database writes or message emits
- Serialized output matches the expected JSON/Protobuf contract
- Error paths (404, 500, timeouts) are tested with realistic failure modes
- Connection pooling and timeouts are exercised in the setup
- Asynchronous background tasks are awaited before assertion
- Migrations are run on the test database before the suite starts
- Environment variables for integration are isolated from local dev
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 78 lines · 34 tokens per session scan A c1b610f0c6b1
integration-testing is a skill published in the GitHub repository kid-sid/codex-spellbook (21 stars, last pushed 3mo ago), licensed MIT. It adds 34 tokens to every session and 629 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
paperclip
Interact with the Paperclip control plane API for task coordination and governance. Use when checking assignments, updating issue status, posting comments, delegating work, managing routines, or calling Paperclip API endpoints.
wayfinder
Plan a huge chunk of work (more than one agent session can hold) as a shared map of decision tickets on your issue tracker, and resolve them one at a time until the way to the destination is clear.
github-labels-query
List GitHub repository labels with perpage pagination and name filtering support.
gsd-audit-milestone
Audit milestone completion against original intent before archiving.
spec-kitty-charter-doctrine
Run charter interview, generation, context, and sync workflows for project governance in Spec Kitty 3.x. Access doctrine artifacts programmatically via DoctrineService. Resolve agent profiles. Load action-scoped governance context iteratively, not all at once. Triggers: "interview for charter", "generate charter"…
bug-triage
Read all open bugs in production/qa/bugs/, re-evaluate priority vs. severity, assign to sprints, surface systemic trends, and produce a triage report. Run at sprint start or when the bug count grows enough to need re-prioritization.