multi-source-data-integrator

multi-source-data-integrator is a skill for Claude Code, Codex from Nero1688/claude-academic-skills. It costs 685 tokens per session (2,765 once invoked), scanned A, original, MIT.

A method for combining independent data sources into one traceable, reproducible research dataset.

In plain words
What is it for?
Use it to link company records across databases, reconcile conflicting financial data, preserve source and retrieval dates, and validate merged datasets.
Why use it?
It addresses mismatched company records, conflicting values, time alignment, and the need to trace every value back to its source.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to link company records across databases, reconcile conflicting financial data, preserve source and retrieval dates, and validate merged datasets.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/nero1688/claude-academic-skills/multi-source-data-integrator
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add Nero1688/claude-academic-skills --skill multi-source-data-integrator
Clone the repo
git clone --depth 1 https://github.com/Nero1688/claude-academic-skills

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for multi-source-data-integrator

README.md
[![agentmods](https://agentmods.dev/badge/skills/nero1688/claude-academic-skills/multi-source-data-integrator/github.svg)](https://agentmods.dev/skills/nero1688/claude-academic-skills/multi-source-data-integrator)
Your own site
<a href="https://agentmods.dev/skills/nero1688/claude-academic-skills/multi-source-data-integrator"><img src="https://agentmods.dev/badge/skills/nero1688/claude-academic-skills/multi-source-data-integrator/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for multi-source-data-integrator

Your own site · 80×15
<a href="https://agentmods.dev/skills/nero1688/claude-academic-skills/multi-source-data-integrator"><img src="https://agentmods.dev/badge/skills/nero1688/claude-academic-skills/multi-source-data-integrator.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 685 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,765 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00685 $0.02765
Opus 5 $0.00342 $0.01383
Sonnet 5 $0.00137 $0.00553
Haiku 4.5 $0.00068 $0.00277

Measured 9d ago against content hash 6fb4171ec31d, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-09, from the pricing page.

Security

Grade A, and why

multi-source-data-integrator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/multi-source-data-integrator/SKILL.md · 95 lines

How it starts

The opening of the file, as written. The whole thing — 95 lines — stays where its author put it; the contents beside it link to each section on GitHub.

多源資料結合方法論架構師(Multi-Source Data Integrator)

核心心法:結合的三種正當理由(先確認你為何要合)

合併前先答「為什麼合這兩個源」,答案決定方法:

  1. 互補涵蓋(complementary):A 有的變數 B 沒有——目的是拼出更完整的變數集。 風險在「對接正確性」。
  2. 互證效度(convergent):A 和 B 都量同一構念——目的是三角驗證。 風險在「把冗餘當獨立證據」。
  3. 事件×序列(event × series):A 給事件日(MOPS 揭露)、B 給報酬/財務序列(TEJ)—— 目的是組事件研究資料。風險在「事件日與序列的時間對齊」。

不同理由→不同嚴謹重點。含糊地「資料越多越好」是審稿人一眼看穿的破綻。

工序 1|實體解析 / 記錄連結(跨源對得上同一家公司)

台灣資料的實體對接陷阱與對策:

  • 主鍵優先序:統一編號(統編)> 公司代號 > 公司名。統編最穩(法人唯一); 股票代號會因轉板(興櫃→上櫃→上市)、更名、暫停/恢復交易、下市而變動或重用。
  • 時間敏感對接:代號在不同期間可能指向不同公司(下市後代號重配)—— 合併 panel 要用「代號×期間」而非只用代號。
  • 名稱模糊比對:只有公司名時,正規化(去「股份有限公司」、全半形、繁簡)後 再模糊比對,且人工複核高風險配對;絕不接受未複核的自動模糊配對進最終資料。
  • 母子公司/合併報表層級:金控與子公司、合併 vs 個體報表,對接前先定研究單位。
  • 產出對接對照表(來源A鍵 ↔ 來源B鍵 ↔ 統編 ↔ 對接方式 ↔ 是否人工複核), 這張表要能附進論文資料附錄。

工序 2|跨源值調解(同一格數字不一致時)

多源對同一「公司-期-變數」給不同值,事先訂規則,不看結果挑:

  • 來源優先序:依變數性質定,並寫進方法節。例:財務數字以查核後財報(MOPS/TEJ 財報)為準;即時事件以 MOPS 重大訊息為準;整齊長序列以 TEJ 為準。
  • 容差區間:數值差在容差內(如四捨五入、單位差)自動取主源;超出容差標記 進衝突清單,人工判讀(常是單位/合併口徑/更正公告造成)。
  • 衝突揭露:報告衝突率與解決方式;衝突率過高(如 >5%)是資料品質警訊, 要回頭查對接是否錯配,而非默默取一個。
  • 絕不「平均兩個源」或「哪個好看取哪個」——這是資料操弄。

工序 3|來源譜系 / lineage(可重現性的命脈)

整合資料集的每一格都要追得回源。最低規格:每個變數欄配一個 _src_asof 伴隨欄 (來源標識 + 擷取/揭露時點)。好處:

  • 審稿人問「這數字哪來的」當場答得出;
  • 免費源端點會變、會更正,_asof 讓你能重建當時快照;
  • 除錯時能定位是哪個源、哪次擷取出問題。 資料附錄要有「來源清單表」:每個源的名稱、版本/擷取日、涵蓋範圍、授權。

工序 4|三角驗證(把多源變成效度證據,而非只是更多資料)

  • 收斂效度:同構念的多源測量該高度相關;相關低要解釋(測量差異 vs 其中一源有誤)。
  • 互補 vs 冗餘的誠實:若兩源高度冗餘,別假裝是兩個獨立證據;若互補,說清各補了什麼。
  • 交叉驗證抽樣:隨機抽 N 筆跨源人工核對(如 30 筆),報一致率——這是資料品質的 直接證據,比任何宣稱都有力。

工序 5|涵蓋與選擇偏誤(合併後樣本的誠實帳)

  • 合併損耗對帳:記錄「A 樣本數 → 內連結後數 → 掉了多少 → 為什麼掉」(對不上/ 單源缺)。這條和 thesis-consistency-audit 的樣本數對帳同源。
  • 選擇偏誤診斷:掉的公司和留的公司比一比(規模/產業/年份分布)——系統性差異 要揭露並討論對外推的影響。
  • 涵蓋差異:各源涵蓋範圍不同(TEJ 上市櫃齊、MOPS 含興櫃、政府資料含未上市), 合併採內連結還是外連結,決定樣本母體,方法節要交代。

Read the full file on GitHub · 95 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 9d ago First seen · 95 lines · 685 tokens per session scan A 6fb4171ec31d

Subscribe to this mod's changes

multi-source-data-integrator is a skill published in the GitHub repository Nero1688/claude-academic-skills (6 stars, last pushed 6d ago), licensed MIT. It adds 685 tokens to every session and 2,765 once invoked, about $0.0034 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

alterlab-pyhealth

Develops, tests, and deploys clinical machine learning models with the PyHealth healthcare AI toolkit. Use when working with electronic health records (EHR), clinical prediction tasks (mortality, readmission, drug recommendation), medical coding systems (ICD, NDC, ATC), physiological signals (EEG, ECG), healthcare…

AlterLab-IEU/AlterLab-Academic-Skills · 117 tokens

alterlab-deepchem

Runs molecular machine learning with DeepChem — diverse featurizers, pre-built MoleculeNet benchmark datasets, and pre-trained models (ChemBERTa, GROVER) for property prediction (ADMET, toxicity, solubility) via traditional ML or graph neural networks. Use when running end-to-end molecular ML experiments that need…

AlterLab-IEU/AlterLab-Academic-Skills · 126 tokens

alterlab-pufferlib

Scales reinforcement learning with PufferLib — high-throughput parallel training (PuffeRL), vectorized environments, and native multi-agent systems achieving 2-10x speedups over standard implementations. Use when scaling RL to millions of steps per second, running vectorized or multi-agent setups, building custom…

AlterLab-IEU/AlterLab-Academic-Skills · 127 tokens

alterlab-shap

Model interpretability and explainability with SHAP (SHapley Additive exPlanations) — feature importance and plots (waterfall, beeswarm, bar, scatter, force, heatmap). Use when explaining ML model predictions, computing feature importance, debugging models, analyzing bias or fairness, comparing models, or implementing…

AlterLab-IEU/AlterLab-Academic-Skills · 115 tokens

alterlab-timesfm

Zero-shot univariate time-series forecasting with Google's TimesFM foundation model, producing point forecasts and prediction intervals from CSV/DataFrame/array inputs, with a preflight system checker for RAM/GPU. Use to forecast any univariate series (sales, sensors, energy, vitals, weather) without training a custom…

AlterLab-IEU/AlterLab-Academic-Skills · 78 tokens

alterlab-esm

Run ESM protein language models — ESM3 for generative multimodal protein design across sequence, structure, and function, and ESM C for efficient embeddings and representations — locally or via the cloud Forge API. Use when working with protein sequences, structures, or function prediction, designing novel proteins…

AlterLab-IEU/AlterLab-Academic-Skills · 89 tokens