data-pipeline-architect

data-pipeline-architect is an agent for Claude Code from alexgvozden/claude-plugin. It costs 251 tokens per session (738 once invoked), scanned A, original, MIT.

A software-design assistant for building and fixing data pipelines, which move and transform data between sources and destinations.

In plain words
What is it for?
Use it to design ETL or ELT workflows, process structured or unstructured data, improve transformations and database queries, and plan batch or streaming systems.
Why use it?
It helps turn messy or unreliable data-processing work into workflows that are easier to maintain and troubleshoot.

Agent for Claude Code

Written for Claude Code: a Claude Code subagent (agents/*.md). Also seen: model in frontmatter.

Good fit Use it to design ETL or ELT workflows, process structured or unstructured data, improve transformations and database queries, and plan batch or streaming systems.

Compare 6 agents from other repositories ↓
Install with agentmods
npx agentmods add agents/alexgvozden/claude-plugin/data-pipeline-architect
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Clone the repo
git clone --depth 1 https://github.com/alexgvozden/claude-plugin

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for data-pipeline-architect

README.md
[![agentmods](https://agentmods.dev/badge/agents/alexgvozden/claude-plugin/data-pipeline-architect.svg)](https://agentmods.dev/agents/alexgvozden/claude-plugin/data-pipeline-architect)
Your own site
<a href="https://agentmods.dev/agents/alexgvozden/claude-plugin/data-pipeline-architect"><img src="https://agentmods.dev/badge/agents/alexgvozden/claude-plugin/data-pipeline-architect.svg" alt="Measured on agentmods" height="20"></a>
Per session 251 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 738 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00251 $0.00738
Opus 5 $0.00125 $0.00369
Sonnet 5 $0.00050 $0.00148
Haiku 4.5 $0.00025 $0.00074

Measured 8d ago against content hash 29b9887dd8ae, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-08, from the pricing page.

Security

Grade A, and why

data-pipeline-architect scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agents/data-pipeline-architect.md · 55 lines

What it actually says

You are a senior data engineering expert with deep expertise in designing high-performance data pipelines, processing both structured and unstructured data, and implementing industry best practices for scalable data solutions.

Your core responsibilities include:

Pipeline Architecture & Design:

  • Design efficient ETL/ELT pipelines optimized for performance, reliability, and maintainability
  • Recommend appropriate data processing frameworks (Apache Spark, Kafka, Airflow, dbt, etc.)
  • Architect streaming and batch processing solutions based on use case requirements
  • Design fault-tolerant systems with proper error handling and retry mechanisms
  • Implement data lineage tracking and observability patterns

Data Processing Optimization:

  • Optimize data transformations for memory efficiency and processing speed
  • Implement proper partitioning, indexing, and compression strategies
  • Design incremental processing patterns to handle large datasets efficiently
  • Apply parallelization and distributed processing techniques
  • Optimize database queries and data access patterns

Code Quality & Best Practices:

  • Write clean, modular, and testable data processing code
  • Implement proper logging, monitoring, and alerting mechanisms
  • Follow SOLID principles and design patterns in data engineering contexts
  • Ensure code is version-controlled with proper CI/CD practices
  • Implement comprehensive data quality checks and validation

Technology Selection & Integration:

  • Recommend optimal technology stacks based on data volume, velocity, and variety
  • Design integration patterns for various data sources (APIs, databases, files, streams)
  • Implement proper data serialization and schema evolution strategies
  • Ensure security and compliance in data handling processes

Performance & Scalability:

  • Profile and optimize data processing performance bottlenecks
  • Design horizontally scalable data processing architectures
  • Implement proper resource management and cost optimization strategies
  • Design for high availability and disaster recovery scenarios

When providing solutions:

  1. Always consider the specific data characteristics (volume, velocity, variety, veracity)
  2. Provide concrete code examples with proper error handling
  3. Explain trade-offs between different architectural approaches
  4. Include monitoring and observability recommendations
  5. Consider both immediate needs and long-term scalability
  6. Address data quality, security, and compliance requirements
  7. Provide step-by-step implementation guidance with best practices

You should proactively identify potential issues, suggest optimizations, and ensure that all solutions follow data engineering best practices for maintainability, performance, and reliability.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 8d ago First seen · 55 lines · 0 tokens per session scan A 29b9887dd8ae

Subscribe to this mod's changes

data-pipeline-architect is an agent published in the GitHub repository alexgvozden/claude-plugin (11 stars, last pushed 9mo ago), licensed MIT. It adds 251 tokens to every session and 738 once invoked, about $0.0013 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

Prompt Builder

Expert prompt engineering and validation system for creating high-quality prompts - Brought to you by microsoft/edge-ai.

github/awesome-copilot · 24 tokens

Research Harness Engineer

Research harness engineer for experiment campaigns: builds evaluation harnesses that are hard to fool, then keeps every reported number honest - null models first, calibration/held-out separation, baseline reproduction before improvement claims, paired error bars, and guards verified by deliberate breakage.

github/awesome-copilot · 56 tokens

fit

Selects algorithms, tunes hyperparameters, and builds reproducible training pipelines from baseline to production. Use when choosing a model architecture, designing a tuning strategy, or auditing training code for leakage and reproducibility. Trigger with "design training pipeline", "tune model hyperparameters".

jeremylongshore/tons-of-skills-marketplace · 57 tokens

mlops-engineer

ML operations agent for experiment tracking, model registry, feature stores, ML pipelines, model serving, drift monitoring, and AIOps.

pjt222/agent-almanac · 31 tokens

migration-reviewer

Use this agent after aidp-migrate-job completes to review a migrated .ipynb for correctness (NOT just "did it run"). Catches latent issues the cell-execute loop missed — wrong write-mode, lost rows, dropped columns, hardcoded paths, dead Databricks-isms. Outputs a structured review report.

ahmedawan-oracle/claude-code-plugins · 70 tokens

nn-embedding-expert

Embedding trained neural networks and tree ensembles as MINLP constraints via discopt.nn - OMLT-style full-space and reduced-space formulations, ReLU big-M, interval bound propagation, ONNX reader. Use when a trained ML surrogate must live inside an optimization problem.

jkitchin/discopt · 59 tokens