Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add ghostinthedata-info/skills --skill performance-tuninggit clone --depth 1 https://github.com/ghostinthedata-info/skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/ghostinthedata-info/skills/performance-tuning)<a href="https://agentmods.dev/skills/ghostinthedata-info/skills/performance-tuning"><img src="https://agentmods.dev/badge/skills/ghostinthedata-info/skills/performance-tuning/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/ghostinthedata-info/skills/performance-tuning"><img src="https://agentmods.dev/badge/skills/ghostinthedata-info/skills/performance-tuning.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00152 | $0.01222 |
| Opus 5 | $0.00076 | $0.00611 |
| Sonnet 5 | $0.00030 | $0.00244 |
| Haiku 4.5 | $0.00015 | $0.00122 |
Grade A, and why
performance-tuning scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 86 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Performance Tuning
Pipeline health is trajectory, not pass/fail. A green DAG that runs 5% slower every month is a ticking clock toward a missed SLA — "slow is harder to see than broken" (Ghost in the Data, "Why Your Pipeline Finishes Later Every Month"). Tune with discipline; don't guess.
Consult docs/agents/platform.md (if present) for which optimisation
features your engine actually has.
The method (in order)
- Measure first — establish a baseline. Instrument each stage (duration AND queue/wait time). Read your orchestrator/transform tool's run artifacts; typically 3–5 models dominate 60–80% of runtime — those are the only ones worth touching. Record real numbers; never eyeball.
- Find the critical path. The longest chain of dependent tasks sets the minimum runtime. Optimising a task not on the critical path changes nothing. A 20-min model running parallel to a 45-min model is not the bottleneck.
- Diagnose with the plan. Pull the query/execution plan or job profile for the dominant stages and read it for full scans, skew, spills, exploding joins, and stale stats. (Engine-specific: see references below.)
- Form a hypothesis, change one thing, re-measure. Bisect against the baseline. If a change didn't move the number, revert it. Re-confirm the critical path after each win — the bottleneck moves.
- Guardrail. Track completion-time drift over weeks; audit dependencies quarterly; alert on a runtime budget, not just on failure.
Common high-leverage fixes (tool-agnostic)
- Partition pruning / clustering. Organise data by the columns you
filter on so queries scan a fraction of the table. Choose keys by how the
table is queried, not how it's loaded; limit to 2–3. Don't wrap the
partition column in a function in the
WHERE— it kills pruning. - Incremental processing. Switch growing fact tables from full rebuild to incremental + idempotent merge so volume growth stops mattering.
- Avoid full scans & view-on-view. Nested views re-execute their whole chain on every query; materialise expensive intermediate layers as tables.
- Filter early, project narrow. Push
WHEREdown; avoidSELECT *. On columnar engines fewer columns = less I/O. - Fix the join. Join on unique keys; match types on both sides; pre-aggregate before joining; broadcast the small side.
- Kill phantom & monster dependencies. Remove refs that no longer reflect
real reads. Keep core spine models lean — push expensive enrichment into a
separate downstream model (
dim_customerslean vsdim_customers_extended) so 40 downstream models don't wait on a propensity score. Rule of thumb: if fewer than half of consumers use an attribute, it doesn't belong on the core model. - Right-size "real-time." Most "we need real-time" is "we need faster batch" — micro-batch every 15–60 min beats new streaming infra. Reserve true streaming for fraud/safety/RTB.
- Event-driven triggers. Replace "schedule at 4am and hope" with trigger-on-arrival to remove safety-margin slack.
What ships with it
4 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 86 lines · 152 tokens per session scan A 7a257e41cdeb
performance-tuning is a skill published in the GitHub repository ghostinthedata-info/skills (5 stars, last pushed 2mo ago), licensed MIT. It adds 152 tokens to every session and 1,222 once invoked, about $0.0008 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
create-pr
Creates a GitHub PR with a Linear-ticket-prefixed title and a decision-led, narrative description for prisma-next. Use when the user wants to create a pull request, open a PR, or submit changes for review.
schema-exploration
Lists tables, describes columns and data types, identifies foreign key relationships, and maps entity relationships in a database. Use when the user asks about database schema, table structure, column types, what tables exist, ERD, foreign keys, or how entities relate.
ha-data-stores
Map of Hope Agent's local data stores and safe read-only query workflow. Use when the user asks where Hope Agent stores data, wants to inspect sessions/messages/memory/logs/background jobs/knowledge indexes/settings, asks the model to query local app data, or debugging requires checking persisted state. Trigger…
supabase
Supabase / PostgREST Row-Level-Security playbook — pull the anon (or leaked servicerole) key out of the frontend JS, map tables from the auto-generated OpenAPI spec, test anonymous RLS READ disclosures (PII/secret leaks), and anonymous RLS WRITE abuse (insert/update/delete — e.g. forging…
nornicdb-cypher-queries
Pick fast, predictable Cypher query shapes in NornicDB — point lookups, batch retrieval, pagination, search, traversal, batched UNWIND/MERGE writes, cleanup, multi-tenant isolation. Use when writing or reviewing Cypher whose latency or throughput matters; maps user intent to the executor's hot-path query templates.
dsql
Build with Aurora DSQL — manage schemas, execute queries, handle migrations, diagnose query plans, diagnose cluster performance, load data, and develop applications with a serverless, distributed SQL database. Covers IAM auth, multi-tenant patterns, MySQL-to-DSQL and PostgreSQL-to-DSQL schema conversion, foreign key…