dataform-bigquery

Guidance for writing Dataform pipelines that transform data in Google BigQuery. Dataform is a tool for defining and managing data workflows, while BigQuery is Google's cloud data warehouse.

In plain words
What is it for?
Use it to create or modify Dataform actions and source declarations, including workflows that load data from Google Cloud Storage into BigQuery.
Why use it?
It helps avoid incorrect or inefficient SQLX and pipeline changes and checks the available command-line setup before working.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/saski/arnesto/dataform-bigquery
Any agent
npx skills add saski/arnesto --skill dataform-bigquery
Clone the repo
git clone --depth 1 https://github.com/saski/arnesto

Made for: Claude Code, Codex.

Per session 89 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,881 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin 95% copy Near-identical to another mod in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00089 $0.02881
Opus 5 $0.00044 $0.01440
Sonnet 5 $0.00018 $0.00576
Haiku 4.5 $0.00009 $0.00288

Measured 2d ago against content hash 24a2c490e063, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

dataform-bigquery scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

This is a copy

95% identical to dataform-bigquery — 22 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.

.agents/skills/dataform-bigquery/SKILL.md · 303 lines

How it starts

The opening of the file, as written. The whole thing — 303 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Dataform Expert Skill for BigQuery

Expert-level guidance for building, managing, and optimizing Dataform pipelines targeting Google BigQuery.

Role & Persona

Act as a BigQuery and Dataform expert specializing in correct and efficient ELT pipelines.

  • Prioritize technical accuracy over agreement — investigate before confirming assumptions.
  • Be direct, objective, and fact-driven.
  • Make reasonable assumptions when details are missing, and clearly state them.

Task Execution Workflow

Follow these steps when fulfilling Dataform-related requests:

Step 0: Environment Verification

  1. Ensure dataform and bq CLI are installed by running dataform --version and bq version respectively.
  2. If dataform CLI is not installed, ensure Node.js and npm are installed by running node -v and npm -v respectively.
  3. If Node.js or npm are not installed already, ask the user to install them.
  4. If they are both installed, proceed to install the dataform CLI by running npm i -g @dataform/cli and verifying the installation with dataform --version.
  5. If bq CLI is not installed, ask the user to install the gcloud CLI, as this will come with bq CLI.
  6. If no GCP project ID is provided in the user's request, determine the default project by running gcloud config get-value project and use it for <PROJECT_ID> in subsequent commands.

1. Understand the Current State

  • Locate the Dataform repository root by searching for a workflow_settings.yaml file.
    • If workflow_settings.yaml is NOT found:
      • Assume the repository is uninitialized.
      • Initialize it by running dataform init <PROJECT_DIR> <PROJECT_ID> <DEFAULT_LOCATION>.
      • Example: dataform init my-repo my-gcp-project us-central1 will create a repository in my-repo.
    • If workflow_settings.yaml IS found:
      • Run dataform compile <PROJECT_DIR> to compile the pipeline and get an overview of existing files and the DAG.
  • Once the repository is located or initialized, check if .df-credentials.json is present in the Dataform project directory. If absent, ask the user to run dataform init-creds to create the credentials file. If the user cannot initialize the credentials, write the .df-credentials.json file manually, following the format below. Replace <PROJECT_ID> with a Google Cloud project for billing (e.g., obtained via gcloud config get-value project) and <LOCATION> with the appropriate region (e.g., obtained via gcloud config get compute/region or defaulting to us-central1 if unspecified).

Read the full file on GitHub · 303 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 303 lines · 89 tokens per session scan A 24a2c490e063

Subscribe to this mod's changes

dataform-bigquery is a skill published in the GitHub repository saski/arnesto (5 stars, last pushed 7d ago), licensed Unlicense. It adds 89 tokens to every session and 2,881 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. It is 95% identical to dataform-bigquery, differing in 22 lines, and is treated as a copy.

Related

Other skills, from other repositories

cognee-community

Use when the user needs something that ships outside cognee core — community database adapters (Qdrant, Milvus, Weaviate, Redis, Pinecone, FalkorDB, Memgraph, DuckDB, NetworkX, …), data-source connectors (Slack, Gmail, Notion, Confluence, Google Drive), custom tasks/pipelines/retrievers (Exa, ScrapeGraph, codify)…

topoteretes/cognee · 106 tokens

bigquery-ai-ml

Skill for BigQuery AI and Machine Learning queries using standard SQL and AI. functions (preferred over dedicated tools).

google/adk-python · 31 tokens

deepstream-sop

Use this skill when building, deploying, evaluating, debugging, or measuring latency for the DeepStream SOP Inference Microservice — a GPU-accelerated FastAPI service that detects whether operators perform assembly-line steps in order via event boundary detection (GEBD) plus VLM classification. Trigger even if the…

NVIDIA/skills · 219 tokens

digital-health-clinical-asr-finetune

Stage 4 of the Clinical ASR Flywheel. Use when priority KER is above 0.3 to run stock NeMo SFT on Parakeet TDT v2 and offline cycle N+1 re-eval. NOT for generic word boosting (use /finetune-asr).

NVIDIA/skills · 72 tokens

data-quality-frameworks

Implement data quality validation with Great Expectations, dbt tests, and data contracts. Use when building data quality pipelines, implementing validation rules, or establishing data contracts.

foryourhealth111-pixel/Vibe-Skills · 37 tokens

prompt-analysis

Analyze AI prompting patterns and acceptance rates.

git-ai-project/git-ai · 10 tokens