Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/datachain-ai/datachain/corenpx skills add datachain-ai/datachain --skill coregit clone --depth 1 https://github.com/datachain-ai/datachainWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00054 | $0.09388 |
| Opus 5 | $0.00027 | $0.04694 |
| Sonnet 5 | $0.00011 | $0.01878 |
| Haiku 4.5 | $0.00005 | $0.00939 |
Grade A, and why
datachain-core scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 776 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are now loaded with expert-level DataChain SDK context. Apply every rule below when generating DataChain Python code.
Scope of this skill
This file is SDK mechanics — how to write DataChain code that runs correctly: API usage, UDF signatures, settings, delta semantics, materialization patterns, saving, exporting.
It does not own methodology. Decisions about which datasets to build, what scope, what shape (Container / Asset / Sense / Task), what fields to save, and when to dialogue with the user about layer choices — those are the CAST methodology, which lives in the datachain-knowledge skill at {knowledge_skill_dir}/CAST.md.
When knowledge is loaded, it is the orchestrator: it plans the layers (CAST §4), invokes the rules in this file to write the code, then runs the KB pipeline. When knowledge is not loaded (raw SDK use, no dc-knowledge/ directory), this file is self-sufficient — CAST doctrine simply does not apply.
If you find yourself reasoning about "should I build a Sense layer here?" or "should this be scoped to the bucket or the directory?" from inside this file, stop — those questions belong upstream. Ask the user to load the knowledge skill, or fall through to a direct solve.
Pre-Generation Checklist
- Every UDF has a known output type. Functions passed to
.map(),.gen(), or.agg()must have their return type resolved. See §2 Rule 2 — the #1 runtime error. - No
from __future__ import annotationsin UDF modules. It stringifies type hints; DataChain's signal-schema resolution then rejects the string-vs-class mismatch. - Bucket access: anonymous or authenticated? Check
dc-knowledge/buckets/for a.mdfile withanon: true/falsein frontmatter. If none, rundatachain bucket status <uri>to detect. Ifdeniedornot found, stop and ask the user. - Heavy-init resources load via
.setup(), not module-level lazy globals:
Lazy globals leak acrosschain.setup(model=lambda: load_model()).map(result=run_model)parallel=Nworkers and hide the dependency from the chain definition. See §2 Rule 20. -
.settings(parallel=N)is the right tool only when the workload benefits. See §2 Rule 6.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 776 lines · 54 tokens per session scan A 6efd3b97c239
datachain-core is a skill published in the GitHub repository datachain-ai/datachain (2,814 stars, last pushed today), licensed Apache-2.0. It adds 54 tokens to every session and 9,388 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
mcp-builder
Guide for creating high-quality MCP (Model Context Protocol) servers that enable LLMs to interact with external services through well-designed tools. Use when building MCP servers to integrate external APIs or services, whether in Python (FastMCP) or Node/TypeScript (MCP SDK).
comet-design
Comet Classic 阶段 2 —— 为 change 产出深度技术 Design Doc。.
comet-hotfix
Comet 预设 —— 通过 open-build-verify-archive 短流程修复已有行为 bug。.
comet-verify
Comet Phase 4: Verify and Close. Invoke with /comet-verify. Verify implementation matches design, handle development branch.
comet-hotfix
Comet preset path: Bug fix / hotfix. Skip brainstorming, directly open → build → verify → archive. Applicable for behavior fixes, scenarios not involving new capability design.
comet-tweak
Comet preset path: Non-bug small changes (tweak). Skip brainstorming and full plan, directly open → lightweight build → light verify → archive. Applicable for copy, configuration, documentation or prompt local optimization.