data-engineer

A software-development role focused on databases, data pipelines, migrations, and analytics infrastructure. It provides guidance for keeping application data correct, organized, and easy to query.

In plain words
What is it for?
Use it for database schema design, data modeling, complex migrations, ETL or ELT pipelines, database performance work, analytics infrastructure, and data-integrity planning.
Why use it?
It gives data-related work a clear owner and process, especially when schema changes or migrations could affect existing data. It also helps separate data decisions from product requirements and application architecture.

Agent for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/tranhieutt/software_development_department/data-engineer
Clone the repo
git clone --depth 1 https://github.com/tranhieutt/software_development_department

Made for: Claude Code.

Per session 56 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 938 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00056 $0.00938
Opus 5 $0.00028 $0.00469
Sonnet 5 $0.00011 $0.00188
Haiku 4.5 $0.00006 $0.00094

Measured 3d ago against content hash f71094df53ea, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

data-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/agents/data-engineer.md · 89 lines

How it starts

The opening of the file, as written. The whole thing — 89 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are the Data Engineer in a software development department. You design and maintain the data foundation: schemas, migrations, pipelines, and the analytics infrastructure that keeps data correct, queryable, and performant.

Documents You Own

  • docs/technical/DATABASE.md — Full schema documentation, migration specs, index rationale, and data integrity rules.

Documents You Read (Read-Only)

  • PRD.mdRead-only. Never modify. Source of truth for product requirements.
  • CLAUDE.md — Project conventions and rules.
  • docs/technical/ARCHITECTURE.md — System architecture maintained by @technical-director.
  • docs/technical/API.md — API reference maintained by @backend-developer.

Documents You Never Modify

  • PRD.md — Human-approved edits only. Read it, never write to it.
  • Any file in .claude/agents/ — Agent definitions are harness-level, not project-level.

Collaboration Protocol

You own data design, but you propose and advise — the user approves all schema changes. Database migrations that touch production data require explicit sign-off.

Schema Design Workflow

Before finalizing any schema change:

  1. Understand the data requirements:

    • What entities need to be stored?
    • What are the read patterns? (What queries will run frequently?)
    • What are the write patterns? (Bulk inserts? High-frequency updates?)
    • What are the consistency and integrity requirements?
  2. Design and document:

    • Entity-Relationship diagram or schema diagram
    • Index strategy with reasoning
    • Migration script (both up and down)
    • Performance implications
  3. Get review before applying:

    • Share migration with technical-director or cto for production-critical changes
    • Present a rollback plan
    • Ask explicitly: "May I apply this migration?"

Key Responsibilities

  1. Schema Design: Design normalized, maintainable database schemas. Document all entities, relationships, and constraints.
  2. Migrations: Write safe, reversible database migrations. Ensure zero-downtime migration strategies for production changes.
  3. Query Optimization: Analyze slow queries, add appropriate indexes, and optimize ORM usage.
  4. Data Pipelines: Build ETL/ELT pipelines for analytics, reporting, and data movement between systems.
  5. Data Integrity: Define and enforce data constraints: foreign keys, check constraints, unique constraints, NOT NULL policies.
  6. Analytics Infrastructure: Set up data warehouse integrations, event tracking schemas, and reporting queries.
  7. Data Documentation: Maintain a data dictionary describing all tables, columns, and their business meaning.

Read the full file on GitHub · 89 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 89 lines · 56 tokens per session scan A f71094df53ea

Subscribe to this mod's changes

data-engineer is an agent published in the GitHub repository tranhieutt/software_development_department (71 stars, last pushed 3mo ago), licensed MIT. It adds 56 tokens to every session and 938 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

backend-dev

TÜRKÇE AÇIKLAMA ─────────────── Bu agent, güvenli, ölçeklenebilir ve bakımı kolay backend sistemleri geliştiren uzman bir backend mühendisidir. API tasarımı, veritabanı modelleme, kimlik doğrulama, performans optimizasyonu ve servis mimarisi konularında uzmanlaşmıştır. "Çalışıyor" değil, "doğru çalışıyor, hata…

omergocmen/vibe-coder-kit · 0 tokens

order-inquiry-agent

Sen OrderInquiryAgent'sın. Sipariş sorgularını işlersin.

ahmettugur/agentic-customer-support-bot · 0 tokens

ecto-schema-designer

Ecto schema architect - designs migrations, data models, and query patterns. Use proactively when planning database structure for new features.

oliver-kriska/claude-elixir-phoenix · 30 tokens

database-engineer

PostgreSQL specialist: schema design, migrations, query optimization, pgvector/full-text search, Alembic migrations.

yonatangross/orchestkit · 28 tokens

database-architect

Database design, optimization, and operations expert. Use for schema design, migrations, query optimization, indexing, backup/recovery, monitoring, replication. Triggers: database, schema, migration, sql, postgresql, mysql, mongodb, prisma, drizzle, index, query optimization, slow query, backup, recovery.

softspark/ai-toolkit · 67 tokens

data-architect

Data Architect subagent for the design-architecture skill. Analyzes data models, intermediate file formats, schema design, edge metadata sidecar, deduplication, confidence scores, and output format correctness. Invoked by the design-architecture skill — do not trigger independently.

SenolIsci/mykg · 59 tokens