exploratory-data-analysis

exploratory-data-analysis is a skill for Claude Code, Codex from thuong-nc/perlytics-skill. It costs 28 tokens per session (1,265 once invoked), scanned A, original, Apache-2.0.

A first-pass examination of a dataset's structure, values, distributions, concentrations, and time patterns. It describes what appears in the data and suggests useful follow-up questions without claiming to explain the causes.

In plain words
What is it for?
Use it to inventory an unfamiliar dataset, find broad patterns, see where data is concentrated, and decide which focused analysis should come next.
Why use it?
It gives you an informed starting point when the business context or analytical direction is unclear. It helps identify notable patterns and anomalies before deeper investigation.

Skill for Claude CodeCodex

Part of the perlytics-skill plugin — 16 skills shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/thuong-nc/perlytics-skill/exploratory-data-analysis
Any agent
npx skills add thuong-nc/perlytics-skill --skill exploratory-data-analysis
Clone the repo
git clone --depth 1 https://github.com/thuong-nc/perlytics-skill

Made for: Claude Code, Codex.

Or install perlytics-skill, the plugin that ships this one along with the rest of its 16 skills.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for exploratory-data-analysis

README.md
[![agentmods](https://agentmods.dev/badge/skills/thuong-nc/perlytics-skill/exploratory-data-analysis.svg)](https://agentmods.dev/skills/thuong-nc/perlytics-skill/exploratory-data-analysis)
Your own site
<a href="https://agentmods.dev/skills/thuong-nc/perlytics-skill/exploratory-data-analysis"><img src="https://agentmods.dev/badge/skills/thuong-nc/perlytics-skill/exploratory-data-analysis.svg" alt="Measured on agentmods" height="20"></a>
Per session 28 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,265 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00028 $0.01265
Opus 5 $0.00014 $0.00633
Sonnet 5 $0.00006 $0.00253
Haiku 4.5 $0.00003 $0.00127

Measured 5d ago against content hash 9105014b24ee, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

exploratory-data-analysis scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/exploratory-data-analysis/SKILL.md · 109 lines

How it starts

The opening of the file, as written. The whole thing — 109 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Exploratory Data Analysis

Purpose

Understand a dataset's structure, distributions, concentrations, and notable patterns before any hypothesis-driven analysis begins.

When to use

Use this skill when:

  • a user shares a dataset with no specific analytical question ("analyze this CSV," "what's in here?")
  • you need to understand what a dataset contains before choosing which analytical skill to apply
  • the business context is unclear and the data itself needs to inform the direction
  • a first look is needed to surface what questions the data is well-positioned to answer

When not to use

Do not use this skill when:

  • the analytical question is already specific and well-framed - proceed directly to the relevant skill
  • data quality has not yet been checked - run data-quality-check first if the dataset is new and untested

Required thinking discipline

  • Describe first, conclude never. EDA surfaces patterns and candidate questions - it does not answer them.
  • Do not force a narrative on the data. If nothing is surprising, say so.
  • Surface concentrations and anomalies without explaining them - explanation requires hypothesis and evidence.
  • End by recommending which analytical direction the data supports, not by stating what the data means.
  • Evidence constraint: Every conclusion must cite specific data — a number, a rate, a segment, or a timeframe. Do not speculate without evidential basis. If data is insufficient, state what is missing rather than asserting an unsupported inference.

Workflow

  1. Inventory: Record the dataset dimensions (row count, column count, time range if a time column exists, primary entity if identifiable). State what each column appears to represent.

  2. Distribution summary:

    • For numeric columns: range, approximate mean and median, skew direction, and whether extreme values exist.
    • For categorical columns: cardinality (how many unique values), top 5 values and their share, and whether there is a long tail.
    • Flag any column where one value dominates more than 80% of rows - this limits the column's usefulness for segmentation.

Read the full file on GitHub · 109 lines

Files

What ships with it

2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 109 lines · 28 tokens per session scan A 9105014b24ee

Subscribe to this mod's changes

exploratory-data-analysis is a skill published in the GitHub repository thuong-nc/perlytics-skill (5 stars, last pushed 4mo ago), licensed Apache-2.0. It adds 28 tokens to every session and 1,265 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

happiness-skill

当用户问「怎么才能更幸福/为什么得到了还不满足/怎么减少焦虑」时调用。 核心理念: 幸福是缺憾感清空的默认状态, 是可训练的技能; 欲望是与自己的契约(得到前不快乐), 同时只留一个重大欲望; 活在当下。 不适用于: 临床抑郁等需要专业治疗的场景(本书方法不能替代医疗)。 Triggers: 幸福/不快乐/欲望/焦虑/知足/活在当下/happiness/desire/anxiety.

kangarooking/cangjie-skill · 136 tokens

docx-comment-reply

Reply to comments (批注) in Word .docx/.doc files: extract comment context, draft replies, write threaded replies back, and validate OOXML.

foryourhealth111-pixel/Vibe-Skills · 39 tokens

sn-image-imitate

Generates a new image that imitates the style of a reference image while updating content based on user intent. Uses a three-stage pipeline: image annotation (long caption), caption rewriting, and image generation. Use when user asks to "imitate style", "保持这个风格重画", "按这张图风格生成", or "style transfer with new content".

OpenSenseNova/SenseNova-Skills · 82 tokens

explaining-machine-learning-models

Explain trained machine learning models through feature attribution, local explanations, and behavior summaries. Use as an explicit/manual helper once a model already exists, not for training ownership, leakage auditing, or general ML strategy selection.

foryourhealth111-pixel/Vibe-Skills · 49 tokens

jobs-to-be-done

Discover what customers truly need by analyzing the "job" they hire your product to do. Use when the user mentions "customer discovery", "why customers churn", "what job does this solve", "competing against luck", "product-market fit", "switching behavior", "milkshake moment", or "functional vs emotional jobs". Also…

wondelai/skills · 137 tokens

refactoring-ui

Audit and fix visual hierarchy, spacing, color, and depth in web UIs. Use when the user mentions "my UI looks off" (or amateur/unprofessional), "fix the design", "Tailwind styling", "color palette", "visual hierarchy", "design system", "spacing scale", or "component styling". Also trigger when building consistent…

wondelai/skills · 132 tokens