data-quality-audit

data-quality-audit is a skill for Claude Code, Codex from antonrisch/db-skills. It costs 46 tokens per session (644 once invoked), scanned A, original, MIT.

A database checking procedure that looks for missing values, duplicate records, broken references, invalid ranges, and related correctness problems. It produces a short findings report and suggested SQL fixes.

In plain words
What is it for?
Use it to investigate data bugs, check migrations or backfills, and verify analytics data against rules such as unique keys, required fields, valid references, and allowed values.
Why use it?
It helps explain why data is inconsistent, such as when reports disagree, migrations change records incorrectly, or rows are missing links to other tables.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to investigate data bugs, check migrations or backfills, and verify analytics data against rules such as unique keys, required fields, valid references, and allowed values.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/antonrisch/db-skills/data-quality-audit
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add antonrisch/db-skills --skill data-quality-audit
Clone the repo
git clone --depth 1 https://github.com/antonrisch/db-skills

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for data-quality-audit

README.md
[![agentmods](https://agentmods.dev/badge/skills/antonrisch/db-skills/data-quality-audit/github.svg)](https://agentmods.dev/skills/antonrisch/db-skills/data-quality-audit)
Your own site
<a href="https://agentmods.dev/skills/antonrisch/db-skills/data-quality-audit"><img src="https://agentmods.dev/badge/skills/antonrisch/db-skills/data-quality-audit/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for data-quality-audit

Your own site · 80×15
<a href="https://agentmods.dev/skills/antonrisch/db-skills/data-quality-audit"><img src="https://agentmods.dev/badge/skills/antonrisch/db-skills/data-quality-audit.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 46 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 644 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00046 $0.00644
Opus 5 $0.00023 $0.00322
Sonnet 5 $0.00009 $0.00129
Haiku 4.5 $0.00005 $0.00064

Measured 8d ago against content hash 5765cd941e34, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-08, from the pricing page.

Security

Grade A, and why

data-quality-audit scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/data-quality-audit/SKILL.md · 133 lines

How it starts

The opening of the file, as written. The whole thing — 133 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Data Quality Audit

Quick Start

Goal: find correctness issues quickly and return a small, actionable report.

When to use this skill

  • bug reports that smell like data drift (missing rows, double counts, weird nulls)
  • after migrations/backfills to prove invariants still hold
  • when analytics numbers disagree between systems

Before you run checks

  • database engine
  • target tables + expected primary keys
  • expected invariants (unique, not null, fk relationships, allowed ranges)
  • performance constraints (large tables, peak hours)

Workflow (default)

  1. pick 1–3 core tables for the issue
  2. run cheap checks first (nulls, duplicates on keys)
  3. run relationship checks (orphans)
  4. run domain checks (ranges, enums)
  5. write a short findings report + safe remediation plan

Core checks (portable sql)

null checks

SELECT COUNT(*) AS null_count
FROM your_table
WHERE important_col IS NULL;

duplicates on a candidate key

SELECT key_col, COUNT(*) AS c
FROM your_table
GROUP BY key_col
HAVING COUNT(*) > 1
ORDER BY c DESC
LIMIT 50;

orphan rows (broken references)

SELECT COUNT(*) AS orphan_count
FROM child c
LEFT JOIN parent p ON p.id = c.parent_id
WHERE c.parent_id IS NOT NULL
  AND p.id IS NULL;

invalid ranges

SELECT COUNT(*) AS bad_count
FROM your_table
WHERE amount < 0;

time sanity (example)

SELECT COUNT(*) AS bad_count
FROM your_table
WHERE created_at > NOW();

Performance-safe tips

  • always start with count(*) + targeted where clauses
  • add limit when inspecting example rows
  • scope by time window if tables are huge (last 7/30 days)
  • prefer indexed predicates (id ranges, created_at) for sampling

Remediation patterns

fix duplicates

  • decide on a canonical row rule (latest by updated_at, highest priority status, etc.)
  • write a deterministic dedupe query
  • add a unique constraint or unique index after cleanup

fix orphans

  • pick policy: delete orphans, reattach to parent, or set fk to null
  • add fk constraint after data is corrected

Read the full file on GitHub · 133 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 8d ago First seen · 133 lines · 46 tokens per session scan A 5765cd941e34

Subscribe to this mod's changes

data-quality-audit is a skill published in the GitHub repository antonrisch/db-skills (9 stars, last pushed 7mo ago), licensed MIT. It adds 46 tokens to every session and 644 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

alloydb-basics

Manages clusters, instances, and backups for AlloyDB for PostgreSQL, and integrates with AlloyDB Model Context Protocol (MCP) tools for automated database operations. Use when creating, configuring, or administering AlloyDB databases. Do NOT use for general PostgreSQL instances (e.g. Cloud SQL) or other GCP databases.

google/skills · 72 tokens

bigquery-bigframes

Generates Python code using BigQuery DataFrames (BigFrames), the pandas/scikit-learn-style API over BigQuery. Use when writing BigFrames code or doing pandas-style dataframe/ML work against BigQuery (e.g. in a notebook). Don't use for SQL-first workflows or the google-cloud-bigquery client library — use…

google/skills · 76 tokens

cloud-databases-onboarding

Guides users through discovering their database requirements, recommends a Google Cloud database based on a recommendation matrix, and assists in database creation. Use when a user asks 'What database service should I use?', 'Help me pick a database', or when a user wants to create a new database on Google Cloud.…

google/skills · 83 tokens

spanner-basics

Assists in provisioning instances and databases, designing performant schemas, and querying data in Spanner. Use when designing primary keys, writing SQL queries or client library code, or diagnosing performance issues.

google/skills · 43 tokens

obsidian-bases

Create and edit Obsidian Bases (.base files) with views, filters, formulas, and summaries. Use when working with .base files, creating database-like views of notes, or when the user mentions Bases, table views, card views, filters, or formulas in Obsidian.

kepano/obsidian-skills · 63 tokens

supabase

Use when doing ANY task involving Supabase. Triggers: Supabase products (Database, Auth, Edge Functions, Realtime, Storage, Vectors, Cron, Queues); client libraries and SSR integrations (supabase-js, @supabase/ssr) in Next.js, React, SvelteKit, Astro, Remix; auth issues (login, logout, sessions, JWT, cookies…

supabase/agent-skills · 185 tokens