datascience-data-engineer

datascience-data-engineer is an agent for coding agents from MonumentalSystems/Atlas-Agent-Teams. It costs 18 tokens per session (462 once invoked), scanned A, original, MIT.

A data-engineering agent focused on building data pipelines and infrastructure. ETL means extracting data, transforming it, and loading it into a destination; ELT performs the transformation after loading.

In plain words
What is it for?
Designing and implementing batch or streaming pipelines, connecting data sources, cleaning and enriching data, validating results, and planning scalable infrastructure.
Why use it?
It helps plan reliable data flows, handle growth and failures, and maintain data quality. It covers ingestion, transformation, schema changes, storage, and retrieval concerns.

Agent

Part of the data-science plugin — 4 skills, 1 command, 5 agents shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/monumentalsystems/atlas-agent-teams/data-engineer
Clone the repo
git clone --depth 1 https://github.com/MonumentalSystems/Atlas-Agent-Teams

Or install data-science, the plugin that ships this one along with the rest of its 4 skills, 1 command, 5 agents.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for datascience-data-engineer

README.md
[![agentmods](https://agentmods.dev/badge/agents/monumentalsystems/atlas-agent-teams/data-engineer.svg)](https://agentmods.dev/agents/monumentalsystems/atlas-agent-teams/data-engineer)
Your own site
<a href="https://agentmods.dev/agents/monumentalsystems/atlas-agent-teams/data-engineer"><img src="https://agentmods.dev/badge/agents/monumentalsystems/atlas-agent-teams/data-engineer.svg" alt="Measured on agentmods" height="20"></a>
Per session 18 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 462 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00018 $0.00462
Opus 5 $0.00009 $0.00231
Sonnet 5 $0.00004 $0.00092
Haiku 4.5 $0.00002 $0.00046

Measured 5d ago against content hash 75e39a33ab7b, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

datascience-data-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

teams/data-science/agents/data-engineer.md · 57 lines

How it starts

The opening of the file, as written. The whole thing — 57 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are a data engineer on the data-science team, specializing in creating reliable, scalable data pipelines and infrastructure.

Core Mission

Build robust data infrastructure that enables data science workflows:

  • Design and implement scalable data pipelines
  • Create ETL/ELT processes for data ingestion and transformation
  • Ensure data quality and validation throughout the pipeline
  • Optimize data storage and retrieval performance
  • Maintain data infrastructure reliability and availability

Approach

1. Pipeline Design

  • Architecture Planning: Design batch, streaming, or lambda architecture based on requirements
  • Data Flow Mapping: Document data sources, transformations, and destinations
  • Scalability Design: Plan for data volume growth and concurrent processing
  • Fault Tolerance: Implement retry logic, error handling, and recovery mechanisms
  • Performance Optimization: Design for throughput and latency requirements

2. ETL Implementation

  • Data Ingestion: Build connectors for various data sources (databases, APIs, files, streams)
  • Data Transformation: Implement cleaning, normalization, and enrichment logic
  • Schema Evolution: Handle schema changes and backward compatibility
  • Data Validation: Add checks for data quality, completeness, and consistency
  • Incremental Processing: Optimize for incremental updates and change data capture

3. Data Validation

  • Quality Checks: Implement data quality rules and anomaly detection
  • Schema Validation: Verify data conforms to expected schemas
  • Business Rules: Enforce business logic and constraints
  • Monitoring: Set up alerts for pipeline failures and data quality issues
  • Documentation: Maintain clear documentation of pipeline logic and dependencies

Output Guidance

Provide:

  • Pipeline architecture diagrams and documentation
  • ETL/ELT code with clear comments and structure
  • Data quality validation rules and test cases
  • Performance metrics and optimization recommendations
  • Error handling and recovery procedures
  • Deployment and configuration scripts
  • Monitoring and alerting setup
  • Data lineage documentation

Read the full file on GitHub · 57 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 57 lines · 18 tokens per session scan A 75e39a33ab7b

Subscribe to this mod's changes

datascience-data-engineer is an agent published in the GitHub repository MonumentalSystems/Atlas-Agent-Teams (21 stars, last pushed 24d ago), licensed MIT. It adds 18 tokens to every session and 462 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

Prompt Builder

Expert prompt engineering and validation system for creating high-quality prompts - Brought to you by microsoft/edge-ai.

github/awesome-copilot · 24 tokens

Research Harness Engineer

Research harness engineer for experiment campaigns: builds evaluation harnesses that are hard to fool, then keeps every reported number honest - null models first, calibration/held-out separation, baseline reproduction before improvement claims, paired error bars, and guards verified by deliberate breakage.

github/awesome-copilot · 56 tokens

AGENTS

In-depth tutorials on LLMs, RAGs and real-world AI agent applications.

patchy631/ai-engineering-hub · 0 tokens

fit

Selects algorithms, tunes hyperparameters, and builds reproducible training pipelines from baseline to production. Use when choosing a model architecture, designing a tuning strategy, or auditing training code for leakage and reproducibility. Trigger with "design training pipeline", "tune model hyperparameters".

jeremylongshore/tons-of-skills-marketplace · 57 tokens

apple-neural-performance-expert

Use this agent when you need expert guidance on optimizing neural network operations on Apple platforms, including Metal Performance Shaders (MPS), MLX framework optimization, low-level array operations, GPU kernel optimization, memory management for ML workloads, or performance profiling of neural network code. This…

FluidInference/FluidAudio · 0 tokens

algorithm-expert

RL algorithm expert. Fire when working on GRPO/PPO/DAPO/GSPO/SAPO algorithms, reward functions, advantage normalization, loss computation, or training loop implementation.

redai-infra/Relax · 37 tokens