developing-incremental-models

A guide for building and troubleshooting dbt incremental models. dbt is a tool that turns SQL into managed data models, while an incremental model updates only new or changed data instead of rebuilding everything.

In plain words
What is it for?
Use it when creating or repairing models that append, merge, upsert, partition data, or handle late-arriving records.
Why use it?
Incremental models can be faster but are harder to design correctly, especially when records change later or unique keys contain duplicates. The guide helps choose an update strategy and check for data drift.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/altimateai/data-engineering-skills/developing-incremental-models
Any agent
npx skills add AltimateAI/data-engineering-skills --skill developing-incremental-models
Clone the repo
git clone --depth 1 https://github.com/AltimateAI/data-engineering-skills

Made for: Claude Code, Codex.

Per session 112 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,232 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00112 $0.02232
Opus 5 $0.00056 $0.01116
Sonnet 5 $0.00022 $0.00446
Haiku 4.5 $0.00011 $0.00223

Measured 2d ago against content hash 0d4c3168f224, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

developing-incremental-models scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

Copies of this mod

1 near-identical copy found in the catalogue:

skills/dbt/developing-incremental-models/SKILL.md · 325 lines

How it starts

The opening of the file, as written. The whole thing — 325 lines — stays where its author put it; the contents beside it link to each section on GitHub.

dbt Incremental Model Development

Choose the right strategy. Design the unique_key carefully. Handle edge cases.

When to Use Incremental

Scenario Recommendation
Source data < 10M rows Use table (simpler, full refresh is fast)
Source data > 10M rows Consider incremental
Source data updated in place Use incremental with merge strategy
Append-only source (logs, events) Use incremental with append strategy
Partitioned warehouse data Use insert_overwrite if supported

Default to table unless you have a clear performance reason for incremental.

Critical Rules

  1. ALWAYS test with --full-refresh first before relying on incremental logic
  2. ALWAYS verify unique_key is truly unique in both source and target
  3. If merge fails 3+ times, check unique_key for duplicates
  4. Run full refresh periodically to prevent data drift

Workflow

1. Confirm Incremental is Needed

# Check source table size
dbt show --inline "select count(*) from {{ source('schema', 'table') }}"

If count < 10 million, consider using table instead. Incremental adds complexity.

2. Understand the Source Data Pattern

Before choosing a strategy, answer:

  • Is data append-only? (new rows added, never updated)
  • Are existing rows updated? (need merge/upsert)
  • Is there a reliable timestamp? (for filtering new data)
  • What's the unique identifier? (for merge matching)
# Check for timestamp column
dbt show --inline "
  select
    min(updated_at) as earliest,
    max(updated_at) as latest,
    count(distinct date(updated_at)) as days_of_data
  from {{ source('schema', 'table') }}
"

3. Choose the Right Strategy

Strategy Use When How It Works
append Data is append-only, no updates INSERT only, no deduplication
merge Data can be updated MERGE/UPSERT by unique_key
delete+insert Data updated in batches DELETE matching rows, then INSERT
insert_overwrite Partitioned tables (BigQuery, Spark) Replace entire partitions

Read the full file on GitHub · 325 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 325 lines · 112 tokens per session scan A 0d4c3168f224

Subscribe to this mod's changes

developing-incremental-models is a skill published in the GitHub repository AltimateAI/data-engineering-skills (122 stars, last pushed 1mo ago), licensed MIT. It adds 112 tokens to every session and 2,232 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

pr-verify

Verify a Docglow change actually works before submitting or merging a PR. Runs the conformance suite, then a behavioral verification pass (flag matrix, artifact-join spot checks, pipeline contract sweep, payload budget). Use when reviewing a PR, self-reviewing a branch before opening a PR, or when asked to "verify…

docglow/docglow · 79 tokens

animation-best-practices

CSS and UI animation patterns for responsive, polished interfaces. Use when implementing hover effects, tooltips, button feedback, transitions, or fixing animation issues like flicker and shakiness.

northgraindata/dbt-doctor · 42 tokens

dbt-doctor

Static analysis and health checks for dbt projects. Use before committing SQL/YAML or when enforcing CI quality gates.

northgraindata/dbt-doctor · 28 tokens

datahub-verified-remediation

Use this skill when a source schema change has broken, or is about to break, a downstream transformation and the user wants a fix they can merge — not a summary. Triggers on: "a column was renamed, fix the dbt model", "this field is gone downstream", "schema drift", "our model still selects the old column", "generate…

Marc-Dvci/praxis-datahub · 140 tokens

Data Pipeline Architect

Design and implement robust data pipelines — ETL/ELT, streaming, batch processing. From architecture to code with Airflow, dbt, Kafka, and modern data stack.

demo112/yunqu-ai-skills · 40 tokens

weekly-report

Generate recurring weekly or monthly analytics reports with period-over-period comparison, anomaly detection, and executive summaries. Use when the user asks for a weekly report, monthly KPI review, recurring metrics snapshot, or needs automated period-over-period diffing. Saves templates for one-command re-runs.

adityawrk/analytics-with-claude-code · 59 tokens