Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add wyattowalsh/agents --skill data-pipeline-architectgit clone --depth 1 https://github.com/wyattowalsh/agentsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/wyattowalsh/agents/data-pipeline-architect)<a href="https://agentmods.dev/skills/wyattowalsh/agents/data-pipeline-architect"><img src="https://agentmods.dev/badge/skills/wyattowalsh/agents/data-pipeline-architect/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/wyattowalsh/agents/data-pipeline-architect"><img src="https://agentmods.dev/badge/skills/wyattowalsh/agents/data-pipeline-architect.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00043 | $0.01842 |
| Opus 5 | $0.00022 | $0.00921 |
| Sonnet 5 | $0.00009 | $0.00368 |
| Haiku 4.5 | $0.00004 | $0.00184 |
Grade A, and why
data-pipeline-architect scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 171 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Data Pipeline Architect
Design resilient data movement and transformation systems for batch, streaming, and hybrid workloads.
Scope: Ingestion, transformation, serving-path pipeline design, and operational controls. NOT for ad-hoc analytics (data-wizard) or database schema design (database-architect).
Canonical Vocabulary
| Term | Definition |
|---|---|
| source | Upstream system producing data |
| sink | Destination system receiving data |
| batch window | Time slice processed as a single scheduled unit |
| watermark | Progress marker used to reason about event-time completeness |
| data contract | Agreement on schema, semantics, freshness, and quality |
| checkpoint | Persisted progress state used for restart and recovery |
| late data | Records arriving after their expected processing window |
| quarantine | Isolated holding area for invalid or suspicious records |
| lineage | Trace from source records through transformations to outputs |
| replay | Reprocessing data from a prior point in time |
Dispatch
| $ARGUMENTS | Mode |
|---|---|
design <pipeline> |
Design a new batch or streaming pipeline |
review <pipeline or architecture> |
Audit an existing pipeline |
operate <issue> |
Improve reliability, monitoring, and recovery |
migrate <change> |
Plan a pipeline migration or re-platform |
contract <dataset> |
Define a producer-consumer data contract |
| Natural language about ingestion, ETL, ELT, or streaming | Auto-detect the closest mode |
| Empty | Show the mode menu with examples |
References
| File | Purpose |
|---|---|
references/decision-matrix.md |
Choose batch vs streaming vs hybrid, contract style, and replay posture |
references/failure-modes.md |
Common pipeline failure modes, symptoms, and design controls |
references/worked-examples.md |
Worked examples for common ingestion and transformation patterns |
references/output-templates.md |
Reusable output shapes for each public mode |
What ships with it
13 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- evals/evals.json 6.1 KB
- references/decision-matrix.md 3.9 KB
- references/failure-modes.md 2.7 KB
- references/output-templates.md 1.7 KB
- references/worked-examples.md 3.1 KB
- scripts/asset_toolkit/__init__.py 303 B runs code
- scripts/asset_toolkit/_shared.py 5.6 KB runs code
- scripts/asset_toolkit/common.py 3.0 KB runs code
- scripts/asset_toolkit/package.py 43 KB runs code
- scripts/asset_toolkit/validate_evals.py 11 KB runs code
- scripts/asset_toolkit/validate_hooks.py 13 KB runs code
- scripts/asset_toolkit/validate_skill.py 3.6 KB runs code
- scripts/check.py 2.2 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 171 lines · 43 tokens per session scan A 3c9c34ddb845
data-pipeline-architect is a skill published in the GitHub repository wyattowalsh/agents (5 stars, last pushed 7d ago), licensed MIT. It adds 43 tokens to every session and 1,842 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
gemini-api-agent-platform
Guides the usage of the Gemini API on Agent Platform with the Google Gen AI SDK for enterprise AI applications. Covers SDK usage (Python, JS/TS, Go, Java, C#), capabilities like Live API, tools, multimedia generation, caching, and batch prediction.
open-source
Documentation reference for writing Python code using the browser-use open-source library. Use this skill whenever the user needs help with Agent, Browser, or Tools configuration, is writing code that imports from browseruse, asks about @sandbox deployment, supported LLM models, Actor API, custom tools, lifecycle…
gemini-api-dev
Use this skill when writing code that calls the Gemini API for text generation, multi-turn chat, multimodal understanding, image generation, video generation, streaming responses, background research tasks, function calling, structured output, or migrating from the old generateContent API. Covers SDK usage and best…
deepstream-sop
Use this skill when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether operators perform assembly-line steps in order via event boundary detection (GEBD) plus VLM classification. Trigger even if the…
azure-search-documents-dotnet
Azure AI Search SDK for .NET (Azure.Search.Documents). Use for building search applications with full-text, vector, semantic, and hybrid search. Covers SearchClient (queries, document CRUD), SearchIndexClient (index management), and SearchIndexerClient (indexers, skillsets). Triggers: "Azure Search .NET"…
azure-search-documents-ts
Build search applications using Azure AI Search SDK for JavaScript (@azure/search-documents). Use when creating/managing indexes, implementing vector/hybrid search, semantic ranking, or building agentic retrieval with knowledge bases.