Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add OKHP3/skillz --skill load-testinggit clone --depth 1 https://github.com/OKHP3/skillzWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/okhp3/skillz/load-testing)<a href="https://agentmods.dev/skills/okhp3/skillz/load-testing"><img src="https://agentmods.dev/badge/skills/okhp3/skillz/load-testing/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/okhp3/skillz/load-testing"><img src="https://agentmods.dev/badge/skills/okhp3/skillz/load-testing.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00077 | $0.00939 |
| Opus 5 | $0.00039 | $0.00469 |
| Sonnet 5 | $0.00015 | $0.00188 |
| Haiku 4.5 | $0.00008 | $0.00094 |
Grade A, and why
load-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 92 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Load testing
A load test answers one of two questions, and confusing them wastes the effort:
- Can it handle X?: a specific, known target. Verification.
- Where does it break, and how?: the ceiling and the failure mode. Discovery.
Discovery is usually more valuable. Knowing you handle 1,000 requests per second tells you less than knowing that at 1,200 the connection pool exhausts and every request hangs for 30 seconds rather than failing fast.
1. Model realistic traffic, not a single endpoint
Hammering one endpoint measures that endpoint. Real systems fail through interaction — a slow report query saturating the pool that the login path needs.
Model:
- The mix: which endpoints, in what proportion, from real traffic data
- The shape: steady, spiky, or diurnal. A ramp reveals different problems than a sudden step
- Think time: real users pause. Zero think time produces an unrealistic connection pattern
- The data distribution: everyone hitting one hot row behaves nothing like a spread of keys. This is a very common cause of misleading results
Done when: the generated traffic resembles what production actually sees.
2. Test something that resembles production
A load test against a laptop tells you about the laptop.
Match, or document the difference: instance sizes, replica counts, database size and data volume, network topology, and — critically — caches in a realistic state. A cold cache and a fully warm one give completely different numbers, and neither may be the steady state.
Done when: you can state how the environment differs from production and how that skews the result.
3. Ramp, and watch for the knee
Do not start at target load. Ramp, and watch for the point where the response curve bends.
Below saturation, latency stays roughly flat as throughput rises. At the knee, latency climbs sharply while throughput stops rising. That point is your real capacity, and it is usually well below the number where errors start.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 92 lines · 77 tokens per session scan A 64f4c9b68927
load-testing is a skill published in the GitHub repository OKHP3/skillz (3 stars, last pushed yesterday), licensed MIT. It adds 77 tokens to every session and 939 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
kelly-agent-eval
Review board (Busabase App-in-Skill) that runs a fixed suite of mock test cases against a baseline vs candidate agent version and surfaces rubric-scored regressions before a release. Use when the user invokes $kelly-agent-eval or /kelly-agent-eval, wants to review agent-version regressions, compare baseline vs…
kelly-app-skill-creator-tests
Build, maintain, and run conformance tests for canonical App-in-Skill projects created by kelly-app-skill-creator. Use when a Kelly app skill needs contract checks, local server smoke tests, responsive browser acceptance, temporary open-source Busabase integration, environment-gated Busabase Cloud OAuth verification…
pentest-api-attacker
Test APIs against OWASP API Security Top 10 including discovery, auth abuse, and protocol-specific checks.
pentest-container-k8s
Test Docker and Kubernetes security controls for RBAC abuse, breakout, and secret exposure.
pentest-remediation-validator
Retest remediated findings, detect regressions, and generate remediation status and certification artifacts.
pentest-vuln-analyzer
Correlate scanner results with CVE and exploit intelligence and prioritize by CVSS and exploitability.