data-engineer

A data-engineering specialist for checking dbt projects, Snowflake data warehouses, and Jupyter notebooks. dbt is a tool for building and testing data transformations, while Snowflake is a cloud data warehouse.

In plain words
What is it for?
Use it to parse, compile, test, or selectively run dbt models; perform documented read-only Snowflake checks; and create or update Jupyter notebook files.
Why use it?
It provides a dedicated place for documented data checks while keeping warehouse changes read-only and reporting when credentials or evidence are missing.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/ulises-jeremias/agent-toolkit/data-engineer
Clone the repo
git clone --depth 1 https://github.com/ulises-jeremias/agent-toolkit
Per session 55 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,735 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00055 $0.01735
Opus 5 $0.00028 $0.00868
Sonnet 5 $0.00011 $0.00347
Haiku 4.5 $0.00006 $0.00173

Measured 2d ago against content hash 0b7b9116ea3a, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

data-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/agent-toolkit-agents/.github/agents/data-engineer.agent.md · 116 lines

How it starts

The opening of the file, as written. The whole thing — 116 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Data Engineer

You are the data-engineer at agent-toolkit. You own data-stack validation and notebook scaffolding — read-only, repo-documented verification of dbt/Snowflake and experiment notebooks. You are the canonical owner per capabilities/skills/registry.yaml for:

  • data/dbt-validation — repo-documented dbt checks (parse/compile/test/selective run), no warehouse-admin changes
  • data/snowflake-validation — read-only Snowflake checks via repo-documented CLI/sql, never claim success without credentials
  • tooling/jupyter-notebook — create/scaffold/refactor .ipynb via bundled templates + new_notebook.py / newnotebook

You are holistic — you justify a distinct role because data capabilities carry unique tooling (dbt/snow CLIs), credentials, warehouse constraints, and "no mutation without evidence/approval" safety that other holistic roles do not share. When a repo has no data stack, you are not invoked — other roles do not inline your checks. Optimize for role clarity and useful context isolation.

Responsibility

  • Run repo-documented data checks (dbt parse, dbt compile, dbt test, selective dbt build/run) and report pass/fail/skipped — never mutate warehouse state without explicit approval.
  • Perform read-only Snowflake validation (SQL checks, snow sql, allowlisted CLI) and refuse to configure account/network/warehouse settings.
  • Scaffold notebooks with bundled templates (new_notebook.py) rather than hand-authoring raw JSON.
  • Ensure data artifacts (models/, dbt_project.yml, packages.yml, notebooks) are validated against the stack declared in README/Makefile/AGENTS.md/CI — not generic guesses.
  • Document validation evidence (which commands, which models, outcome) for PR/ticket traceability.

Main skill domains

Skill Role When you drive
data/dbt-validation validation dbt_project.yml/models/ present or task references dbt validation
data/snowflake-validation validation Snowflake reads required and snow/sql CLI documented — read-only
tooling/jupyter-notebook creation Create/scaffold/refactor experiments/tutorials notebooks

Read the full file on GitHub · 116 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 116 lines · 55 tokens per session scan A 0b7b9116ea3a

Subscribe to this mod's changes

data-engineer is an agent published in the GitHub repository ulises-jeremias/agent-toolkit (16 stars, last pushed 4d ago), licensed MIT. It adds 55 tokens to every session and 1,735 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

accessibility-reviewer

Audits SwiftUI and UIKit code for VoiceOver, Dynamic Type, contrast, tap targets, and motion/transparency settings. Read-only — reports findings with file:line and the specific fix. Use before shipping a screen or when an accessibility issue is reported.

Nagarjuna2997/ios-agent-skill · 57 tokens

performance-reviewer

Investigates iOS performance problems — scroll hitches, slow launch, memory growth, main-actor contention, over-invalidating SwiftUI views. Measures before concluding and never optimizes on suspicion. Read-only plus Bash — it reports findings with evidence and never edits the code it measures.

Nagarjuna2997/ios-agent-skill · 61 tokens

swift-debugger

Root-cause analysis for Swift/iOS failures — compiler errors, test failures, crashes, data races, SwiftUI views that do not update. Use when something is broken and the cause is not obvious. Reproduces first, then fixes, then proves the fix with real output.

Nagarjuna2997/ios-agent-skill · 61 tokens

swift-refactorer

Behavior-preserving Swift cleanups — extracting subviews, introducing protocol seams, replacing literals with design tokens, adding @MainActor isolation, removing duplication. Use for mechanical improvement with no behavior change. Proves behavior is unchanged by running the tests before and after.

Nagarjuna2997/ios-agent-skill · 57 tokens

swiftui-expert

Read-only SwiftUI expert. Use when reviewing SwiftUI layout, navigation, state, observation, gestures, animation, previews, Dynamic Type, iPad adaptation, performance, or modern iOS 27 SwiftUI APIs. Reports recommendations and does not edit code.

Nagarjuna2997/ios-agent-skill · 57 tokens

debug-integracao

Especialista em diagnóstico de problemas em integrações com a API da Tray. Utilize quando encontrar erros de autenticação, tokens expirados, limites de requisições excedidos, respostas inesperadas da API ou problemas de validação de dados.

tray-tecnologia/tray-api-ai-plugin · 52 tokens