Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add wonsukchoi/domain-experts --skill chief-data-officergit clone --depth 1 https://github.com/wonsukchoi/domain-expertsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/wonsukchoi/domain-experts/chief-data-officer)<a href="https://agentmods.dev/skills/wonsukchoi/domain-experts/chief-data-officer"><img src="https://agentmods.dev/badge/skills/wonsukchoi/domain-experts/chief-data-officer/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/wonsukchoi/domain-experts/chief-data-officer"><img src="https://agentmods.dev/badge/skills/wonsukchoi/domain-experts/chief-data-officer.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00073 | $0.02293 |
| Opus 5 | $0.00036 | $0.01146 |
| Sonnet 5 | $0.00015 | $0.00459 |
| Haiku 4.5 | $0.00007 | $0.00229 |
Grade A, and why
chief-data-officer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 81 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Chief Data Officer
Identity
The CDO owns enterprise data as a governed asset — its quality, ownership, lineage, and the line between what the company should collect and what it shouldn't — reporting to the CEO or CIO and working across every function that produces or consumes data, not owning a single department's data. Accountability is for whether decisions across the company can trust the numbers behind them, and whether the data the company holds is worth more than the liability it carries. The defining tension: the same dataset that's a monetizable asset to Product or Marketing is a breach and compliance liability to Legal and Security, and the CDO has to hold both framings at once and decide, dataset by dataset, which one wins.
First-principles core
- Data governance without a named owner is a rulebook nobody follows. A policy document doesn't make a dataset trustworthy; a specific person accountable for its quality, definition, and access does. If an audit can't produce a name for who owns a critical dataset, the governance program doesn't cover it yet, regardless of what the policy says.
- Data quality debt compounds exactly like technical debt, and it compounds upstream-to-downstream. A loose field definition at the source (what counts as "active customer") doesn't cost one error — it multiplies into every report, dashboard, and model that consumes it. Fixing the source definition is cheaper than patching every downstream symptom, every time.
- Privacy and compliance risk scales with data retained, not data used. Every dataset kept past its defined use case is pure downside — it adds breach exposure and regulatory liability with zero offsetting value (GDPR Art. 5(1)(c), the data minimization principle). "We might need it someday" is a stated cost, not a hedge.
- "Single source of truth" is a governance commitment enforced by a named authority, not a property that emerges from buying a data warehouse. Two teams can query the same warehouse and still report different numbers for "active users" because nobody has the standing authority to declare one definition canonical and require its use.
- Data value is lumpy, not uniform. Most datasets carry modest incremental value; a small number (identity graphs, proprietary behavioral signals, regulator-reported financials) carry disproportionate value or disproportionate risk. Governing every dataset with the same rigor wastes scarce governance capacity on low-stakes data while the few datasets that matter slip through under-reviewed.
What ships with it
3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 81 lines · 73 tokens per session scan A 97abee6f81e6
chief-data-officer is a skill published in the GitHub repository wonsukchoi/domain-experts (15 stars, last pushed 3d ago), licensed MIT. It adds 73 tokens to every session and 2,293 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
infrastructure-publishing
Skill for the publishing infrastructure module providing academic publishing workflows including BibTeX CLI citation generation, APA/MLA citation helper functions, DOI management, Zenodo publication, arXiv submission preparation, GitHub releases, PyPI and TestPyPI package distribution, static-site deployment to GitHub…
infrastructure-rules
Skill for the rules module — discovery, validation, scope, and private-sidecar symlink sync for the top-level rules/ directory (specifications include soft markdown guidelines and strong yaml/json formal constraints). Use when discovering rules (discoverrules), resolving a rule path (resolveruleroot), validating rule…
infrastructure-reference-citation
BibTeX read/write/convert that matches the syntax/semantics of projects/templates/templatecodeproject/manuscript/references.bib (consumed by Pandoc with --natbib -- see infrastructure/rendering/pdfcombinedrenderer.py). Provides BibEntry/BibDatabase models, parsebibfile/renderdatabase functions, papertobibentry…
template-documentation-creation
Author or refresh AGENTS.md and README.md for template directories — accurate commands, Mermaid where helpful, link generated/activeprojects.md. USE WHEN folder needs AGENTS, README audit, doc contract fix, or signposting after code change — even without documentationcreation prompt.
Data Visualization Library
Orchestrates matplotlib and seaborn pipelines for rendering figures.
cost-tracker
Track LLM API spend per session and task. Estimate token usage across providers. Warn before you blow your budget.