metadata-harmonization-starter

metadata-harmonization-starter is a skill for Claude Code, Codex from ma-compbio-lab/SkillFoundry. It costs 0 tokens per session (313 once invoked), scanned A, original, Apache-2.0.

A tool for combining small metadata tables that use different column names or category labels into one consistent tab-separated file (TSV). It also creates a short JSON summary of the changes.

In plain words
What is it for?
Use it to map source columns to standard fields, normalize values such as sex or condition, and produce consistent metadata outputs.
Why use it?
It removes the manual work of matching equivalent fields and labels across test files before validating or converting the data.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/ma-compbio-lab/skillfoundry/metadata-harmonization-starter
Any agent
npx skills add ma-compbio-lab/SkillFoundry --skill metadata-harmonization-starter
Clone the repo
git clone --depth 1 https://github.com/ma-compbio-lab/SkillFoundry

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for metadata-harmonization-starter

README.md
[![agentmods](https://agentmods.dev/badge/skills/ma-compbio-lab/skillfoundry/metadata-harmonization-starter.svg)](https://agentmods.dev/skills/ma-compbio-lab/skillfoundry/metadata-harmonization-starter)
Your own site
<a href="https://agentmods.dev/skills/ma-compbio-lab/skillfoundry/metadata-harmonization-starter"><img src="https://agentmods.dev/badge/skills/ma-compbio-lab/skillfoundry/metadata-harmonization-starter.svg" alt="Measured on agentmods" height="20"></a>
Per session 0 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 313 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00000 $0.00313
Opus 5 $0.00000 $0.00156
Sonnet 5 $0.00000 $0.00063
Haiku 4.5 $0.00000 $0.00031

Measured 5d ago against content hash c14368ce34ce, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

metadata-harmonization-starter scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

The scan reads SKILL.md. This mod also ships 2 executable files (scripts/run_metadata_harmonization.py, tests/test_run_metadata_harmonization.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/data-acquisition-and-dataset-handling/metadata-harmonization-starter/SKILL.md · 33 lines

What it actually says

Metadata Harmonization Starter

Use this skill to harmonize small metadata tables with inconsistent column names and categorical labels into one canonical TSV plus a compact JSON summary.

What This Skill Does

  • reads one or more tabular metadata files
  • applies a JSON mapping from source columns to canonical fields
  • normalizes selected categorical values such as sex and condition
  • writes a harmonized TSV and a machine-readable summary

When To Use It

  • when you need a starter for metadata-harmonization
  • when multiple small test fixtures use different metadata headers
  • when you want deterministic harmonized outputs before validation or format conversion

Run

python3 skills/data-acquisition-and-dataset-handling/metadata-harmonization-starter/scripts/run_metadata_harmonization.py \
  --input skills/data-acquisition-and-dataset-handling/metadata-harmonization-starter/examples/cohort_a.tsv \
  --input skills/data-acquisition-and-dataset-handling/metadata-harmonization-starter/examples/cohort_b.tsv \
  --mapping skills/data-acquisition-and-dataset-handling/metadata-harmonization-starter/examples/column_mapping.json \
  --out-tsv scratch/metadata-harmonization/harmonized_metadata.tsv \
  --summary-out scratch/metadata-harmonization/harmonized_metadata_summary.json

Notes

  • The starter keeps the mapping external so the same script can be reused for other tiny fixtures.
  • Canonical rows are sorted by sample_id to keep committed outputs deterministic.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 33 lines · 0 tokens per session scan A c14368ce34ce

Subscribe to this mod's changes

metadata-harmonization-starter is a skill published in the GitHub repository ma-compbio-lab/SkillFoundry (38 stars, last pushed 4mo ago), licensed Apache-2.0. It costs nothing until one of its globs matches a file; then it loads 313 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

grounded-review

Review a research report draft with a structured scoring rubric, run a bounded repair loop when needed, and produce the final deliverable report.

gaotiexinqu/OneResearchClaw · 31 tokens

one-report

Run the full One-Report pipeline from one input file through grounding, research, evidence-rich report drafting, final review quality-gating, and export, while reusing existing skills, preserving current contracts, and enforcing strict downstream skill fidelity for every grounded unit.

gaotiexinqu/OneResearchClaw · 54 tokens

grounded-research-lit

Run focused literature and web research from a grounded note. Use when a grounded note already exists and you want targeted research results, opened-link evidence, deeper per-paper analysis materials, optional downloaded literature, and a two-stage literature output (litinitial.md then refined lit.md).

gaotiexinqu/OneResearchClaw · 62 tokens

grounded-summary

Create a rich, evidence-preserving research report draft from a grounded note and its follow-up literature result. This is the main report-writing stage of the middle pipeline, not a compression memo.

gaotiexinqu/OneResearchClaw · 42 tokens

remote-input

Download remote content (arxiv papers, YouTube videos, Bilibili videos) to local storage and route to downstream grounding pipeline. Use when user provides a URL instead of a local file path.

gaotiexinqu/OneResearchClaw · 43 tokens

skill-evolve

Optional sidecar skill for controlled feedback-driven skill evolution. Not part of the default pipeline. Only activates when explicitly requested.

gaotiexinqu/OneResearchClaw · 28 tokens