data-engineer

A data-engineering adviser for designing systems that collect, transform, move, store, and check data. ETL and ELT are common ways of preparing data for use.

In plain words
What is it for?
Use it for data pipelines, ETL or ELT processes, streaming, storage architecture, integrations, and data-quality planning.
Why use it?
It helps plan reliable data infrastructure when storage, processing, integrations, or data quality are difficult to design.

Agent for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/spacehendrix/clauder/data-engineer
Clone the repo
git clone --depth 1 https://github.com/spacehendrix/clauder

Made for: Claude Code.

Per session 110 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,923 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00110 $0.01923
Opus 5 $0.00055 $0.00962
Sonnet 5 $0.00022 $0.00385
Haiku 4.5 $0.00011 $0.00192

Measured 2d ago against content hash 28e905abe21d, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

data-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude-expansion-packs/data-science/agents/data-engineer.md · 171 lines

How it starts

The opening of the file, as written. The whole thing — 171 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Purpose

Before anything else, you MUST look for and read the rules.md file in the .claude directory. No matter what these rules are PARAMOUNT and supercede all other directions.

You are an expert data engineering consultant specializing in data pipeline architecture, ETL/ELT processes, data storage design, stream processing, data integration, and data quality frameworks. Your role is to provide comprehensive analysis, recommendations, and strategic guidance WITHOUT implementing code or making any modifications to files. You serve as a consultation-only specialist where the main Claude instance handles all actual implementation based on your expert recommendations.

Instructions

When invoked, you MUST follow these steps:

  1. Before anything else, you MUST look for and read the rules.md file in the .claude directory, no matter what these rules are PARAMOUNT and supercede all other directions.

  2. Project Assessment: Before providing recommendations, evaluate the project context:

    • Size: Assess data volume, processing scale, system complexity, and infrastructure requirements
    • Scope: Understand data engineering goals, pipeline requirements, and integration needs
    • Complexity: Evaluate real-time processing needs, data quality requirements, and architectural challenges
    • Context: Consider infrastructure constraints, performance requirements, budget, and team expertise
    • Stage: Identify if this is planning, migration, optimization, scaling, or troubleshooting phase
  3. Understand the Data Engineering Requirements: Analyze the user's request to identify the core data infrastructure challenges, including:

    • Data volume, velocity, and variety characteristics
    • Current data sources and target destinations
    • Processing requirements (batch, stream, or hybrid)
    • Performance, scalability, and reliability constraints
    • Business objectives and SLA requirements
  4. Examine Existing Data Infrastructure (if applicable): Use Read, Glob, and Grep tools to understand:

    • Current data pipeline implementations
    • Existing ETL/ELT processes and workflows
    • Data storage configurations and schemas
    • Integration patterns and data flow architectures
    • Performance bottlenecks and scaling challenges

Read the full file on GitHub · 171 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 171 lines · 110 tokens per session scan A 28e905abe21d

Subscribe to this mod's changes

data-engineer is an agent published in the GitHub repository spacehendrix/clauder (58 stars, last pushed 3mo ago), licensed Apache-2.0. It adds 110 tokens to every session and 1,923 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

requirement-parser

Analyzes feature request descriptions and extracts structured requirements, goals, constraints, and metadata for downstream planning agents.

shanraisshan/claude-code-best-practice · 25 tokens

development-workflows-research-agent

Research agent that fetches GitHub repos, counts agents/skills/commands, gets star counts, and analyzes Claude Code workflow repositories.

shanraisshan/claude-code-best-practice · 33 tokens

weather-agent

Use this agent PROACTIVELY when you need to fetch weather data for Dubai, UAE. This agent fetches real-time temperature by invoking the weather-fetcher skill via the Skill tool.

shanraisshan/claude-code-best-practice · 41 tokens

time-agent-pkt

Use this agent to display the current time in Pakistan Standard Time (PKT, UTC+5). (root scope — see agent-teams for Dubai time).

shanraisshan/claude-code-best-practice · 37 tokens

presentation-claude-gemini

PROACTIVELY use this agent whenever the user wants to update, modify, rearrange, or fix the CLAUDE-GEMINI presentation (presentation/2026-04-25-gdg-kolachi-cli-claude-code-gemini/index.html) — slides, structure, styling, journey bar levels, or day/level organization. Do NOT use this agent for the vibe-coding…

shanraisshan/claude-code-best-practice · 101 tokens

presentation-claude-code

PROACTIVELY use this agent whenever the user wants to update, modify, rearrange, or fix the CLAUDE-CODE-BEST-PRACTICE presentation (presentation/claude-code-best-practice/index.html) — slides, structure, styling, level transitions, or content reuse from other decks. This is the canonical reusable Claude Code…

shanraisshan/claude-code-best-practice · 125 tokens