data-engineering

Guidance for moving, transforming, and checking data between systems. It covers pipelines, which are repeatable flows that take data from a source to a destination.

In plain words
What is it for?
Use it to build ETL or ELT jobs, migrate data, design warehouse tables, monitor data quality, capture database changes, or handle changing data formats.
Why use it?
It helps prevent missing, corrupted, or incompatible data from quietly reaching reports, warehouses, or machine-learning systems.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/borhen68/skillengine/data-engineering
Any agent
npx skills add borhen68/SkillEngine --skill data-engineering
Clone the repo
git clone --depth 1 https://github.com/borhen68/SkillEngine

Made for: Claude Code, Codex.

Per session 59 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,734 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00059 $0.02734
Opus 5 $0.00030 $0.01367
Sonnet 5 $0.00012 $0.00547
Haiku 4.5 $0.00006 $0.00273

Measured yesterday against content hash b368cb7aab39, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

data-engineering scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/data-engineering/SKILL.md · 355 lines

How it starts

The opening of the file, as written. The whole thing — 355 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Data Engineering

Overview

Data engineering is the foundation of every data-driven decision. Bad data pipelines silently corrupt analytics, break ML models, and lead to business decisions based on false premises. This skill covers designing pipelines that are correct, observable, and resilient — from ingestion to serving.

The data engineering contract: Every pipeline must guarantee that what lands in the destination is what the source intended, or it must fail loudly. Silent data corruption is the worst failure mode.

When to Use

  • Building ETL, ELT, or streaming data pipelines
  • Designing data warehouse schemas (star, snowflake, data vault)
  • Migrating data between systems or formats
  • Setting up data quality monitoring and anomaly detection
  • Creating CDC (change data capture) pipelines
  • Building feature stores for machine learning
  • Handling schema evolution without breaking consumers

NOT for:

  • Simple one-off data exports (use a script)
  • Real-time systems with sub-second latency requirements (use stream processing)
  • Data science / analysis work (this skill is about moving and transforming data, not interpreting it)

The Data Pipeline Process

Step 1: Define the Data Contract

Before writing any pipeline code, define what correctness means:

DATA CONTRACT:
├── Source: [system, table, API, file format]
├── Destination: [system, table, format]
├── Schema: [field names, types, nullability, defaults]
├── Volume: [records/day, peak throughput, growth rate]
├── Latency: [batch hourly / batch daily / streaming / near-real-time]
├── Quality rules: [uniqueness, referential integrity, range checks]
├── Retention: [how long to keep, compliance requirements]
└── SLA: [acceptable downtime, max lag, error rate threshold]

Schema definition example:

# data_contract.yaml
source:
  system: production_postgres
  table: orders
  
destination:
  system: snowflake
  schema: analytics
  table: fact_orders
  
schema:
  order_id: { type: BIGINT, nullable: false, unique: true }
  customer_id: { type: BIGINT, nullable: false }
  order_date: { type: TIMESTAMP, nullable: false }
  amount: { type: DECIMAL(10,2), nullable: false, min: 0 }
  status: { type: VARCHAR(20), nullable: false, enum: [pending, paid, shipped, cancelled] }
  
quality_rules:
  - column: order_id
    check: not_null
  - column: amount
    check: range
    min: 0
    max: 100000
  - table: fact_orders
    check: referential_integrity
    references: dim_customers.customer_id

Read the full file on GitHub · 355 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 355 lines · 59 tokens per session scan A b368cb7aab39

Subscribe to this mod's changes

data-engineering is a skill published in the GitHub repository borhen68/SkillEngine (17 stars, last pushed 2mo ago), licensed MIT. It adds 59 tokens to every session and 2,734 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

code-review-and-quality

执行多维度代码审查。用于合并任何变更之前;用于审查自己、其他 agent 或人类编写的代码;用于在代码进入主分支前从多个维度评估代码质量。.

vinvcn/addyosmani-agent-skills-zh · 53 tokens

code-simplification

为清晰度简化代码。用于在不改变行为的前提下重构代码以提升清晰度;用于代码能运行但比应有状态更难阅读、维护或扩展时;用于审查已累积不必要复杂度的代码时。.

vinvcn/addyosmani-agent-skills-zh · 64 tokens

doubt-driven-development

在每个非平凡决策成立前,用全新上下文进行对抗式审查。当正确性比速度更重要、处理不熟悉代码、风险较高(生产、安全敏感逻辑、不可逆操作),或任何自信输出现在验证比之后调试更便宜时使用。.

vinvcn/addyosmani-agent-skills-zh · 72 tokens

test-driven-development

用测试驱动开发。用于实现任何逻辑、修复任何 bug,或改变任何行为。用于需要证明代码能工作、收到 bug 报告,或即将修改现有功能时。.

vinvcn/addyosmani-agent-skills-zh · 48 tokens

api-and-interface-design

指导稳定的 API 和接口设计。设计 API、模块边界或任何公共接口时使用。创建 REST 或 GraphQL endpoint、定义模块之间的类型契约,或建立前后端边界时使用。.

vinvcn/addyosmani-agent-skills-zh · 51 tokens

ci-cd-and-automation

自动化 CI/CD pipeline 设置。用于设置或修改构建和部署 pipeline 时;用于需要自动化质量门禁、在 CI 中配置 test runners,或建立部署策略时。.

vinvcn/addyosmani-agent-skills-zh · 46 tokens