codexkit-data-quality-auditor

A review of how trustworthy a dataset is. It checks completeness, accuracy, consistency, timeliness, validity, and uniqueness, then scores the results and records problems.

In plain words
What is it for?
Use it to profile a new table or file, assess a data source, log quality issues, and plan fixes or monitoring rules for analytics and dashboards.
Why use it?
It finds missing, incorrect, stale, duplicated, or conflicting data before those issues produce misleading reports or break a migration or integration.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/hoavdc/codexkit/codexkit-data-quality-auditor
Any agent
npx skills add hoavdc/CodexKit --skill codexkit-data-quality-auditor
Clone the repo
git clone --depth 1 https://github.com/hoavdc/CodexKit

Made for: Claude Code, Codex.

Per session 69 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,173 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00069 $0.01173
Opus 5 $0.00034 $0.00587
Sonnet 5 $0.00014 $0.00235
Haiku 4.5 $0.00007 $0.00117

Measured 2d ago against content hash 89933a5a0837, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

codexkit-data-quality-auditor scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/codexkit-data-quality-auditor/SKILL.md · 143 lines

How it starts

The opening of the file, as written. The whole thing — 143 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Data Quality Auditor

When to Use

  • Before building analytics or dashboards on a new data source
  • When data issues cause downstream report errors
  • During data migration or system integration
  • When establishing data quality monitoring rules

Procedure

Step 1 — Scope & Profiling

Identify the dataset and context:

  • Table/file name, row count, column count
  • Business purpose: what decisions does this data support?
  • Data owner: who is accountable?

Profile the data:

  • Column types (string, numeric, date, boolean, null)
  • Null rates per column
  • Distinct value counts
  • Min/Max/Mean for numerics
  • Date range coverage

Step 2 — Six-Dimension Assessment

Score each dimension 1–5:

Dimension Question Score Method
Completeness Are all required fields populated? % non-null for required fields
Accuracy Do values represent reality? Sample validation against source
Consistency Do related fields agree? Cross-field rule checks
Timeliness Is data current enough for its purpose? Max staleness vs requirement
Validity Do values conform to allowed ranges/formats? Format + range validation
Uniqueness Are there unwanted duplicates? Duplicate rate on key columns

Step 3 — Issue Log

For each issue found:

# Dimension Column(s) Issue Description Severity Records Affected Example
1 Completeness email 12% null in required field High 1,200 row 45: null

Severity: Critical (blocks use) / High (degrades quality) / Medium (cosmetic) / Low (nice-to-fix)

Step 4 — Remediation Plan

For each Critical/High issue:

Issue Root Cause Fix Owner Deadline
Missing emails Optional in old form Backfill from CRM Data Eng Sprint 4

Step 5 — Monitoring Rules

Define ongoing quality checks:

  • Automated checks to run on each data load
  • Alerting thresholds (if quality drops below X%, alert owner)
  • Review cadence (weekly, monthly)

Read the full file on GitHub · 143 lines

Files

What ships with it

4 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 143 lines · 69 tokens per session scan A 89933a5a0837

Subscribe to this mod's changes

codexkit-data-quality-auditor is a skill published in the GitHub repository hoavdc/CodexKit (21 stars, last pushed 3mo ago), licensed MIT. It adds 69 tokens to every session and 1,173 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

reducing-aigc-detection

Systematically reduce AIGC detection rates in academic papers (Chinese/English). Analyzes detection reports, identifies high-impact sections, applies multi-layer rewriting strategies preserving formatting/footnotes, and verifies results. Supports 维普/知网/Turnitin platforms.

telagod/code-abyss · 62 tokens

defending-applications

Application security defense knowledge for builders. Covers Web/API/GraphQL hardening (XSS/SQLi/SSRF/IDOR/BOLA/Mass Assignment/deserialization/upload/path traversal), authentication/authorization (OAuth 2.0/OIDC/JWT/Session/Cookie/SAML/SSO), and LLM application security (prompt injection, jailbreak, RAG poisoning…

telagod/code-abyss · 152 tokens

verification-loop

Evidence-before-assertions workflow. Use before claiming work is done, before release, and after any behavior change in scripts/skills/MCP.

rexleimo/aios · 32 tokens

analyzing-security

Scans code for security vulnerabilities, detects dangerous patterns, and ensures security decisions are documented. Use when running security scans, auditing code, or checking for OWASP issues, injection risks, or sensitive data leaks. Automatically triggered on new modules, security-related changes, or post-refactor.

telagod/code-abyss · 62 tokens

architecting-security

安全架构与治理:威胁建模 (STRIDE/PASTA/LINDDUN)、零信任身份架构、IAM/SSO/MFA/PAM、合规框架 (SOC2/PCI/HIPAA/GDPR)、DLP、隐私工程、安全控制设计。Use when designing security architecture, threat modeling new systems, implementing zero-trust identity, designing IAM/SSO/PAM, building compliance evidence chains, or planning privacy-by-design.

telagod/code-abyss · 104 tokens

building-agent-systems

AI agent and LLM system engineering reference covering single-agent dev (ReAct, tool calling, plan-execute), multi-agent coordination (swarm, role decomposition, file locking), LLM security (prompt injection, jailbreak defense, output filtering), RAG architecture (chunking, hybrid retrieval, rerank), and prompt…

telagod/code-abyss · 110 tokens