Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add CUHK-AIM-Group/NeuroClaw --skill harmonization-toolgit clone --depth 1 https://github.com/CUHK-AIM-Group/NeuroClawWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/cuhk-aim-group/neuroclaw/harmonization-tool)<a href="https://agentmods.dev/skills/cuhk-aim-group/neuroclaw/harmonization-tool"><img src="https://agentmods.dev/badge/skills/cuhk-aim-group/neuroclaw/harmonization-tool/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/cuhk-aim-group/neuroclaw/harmonization-tool"><img src="https://agentmods.dev/badge/skills/cuhk-aim-group/neuroclaw/harmonization-tool.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00119 | $0.03869 |
| Opus 5 | $0.00060 | $0.01935 |
| Sonnet 5 | $0.00024 | $0.00774 |
| Haiku 4.5 | $0.00012 | $0.00387 |
Grade A, and why
harmonization-tool scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 275 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Harmonization Tool Skill (Cross-Site Feature Alignment Layer)
Overview
harmonization-tool is the NeuroClaw cross-cutting layer that sits between dataset skills (ABIDE, ADHD-200, ABCD, HCP, UKB, ...) and model skills (BrainGNN, BNT, IBGNN, LGGNN, BrainNetCNN, FM-APP, SVM, SpaceNet, ...).
Its job: take subject-level features extracted by dataset skills and remove technical / batch variance introduced by site, scanner, field strength, sequence, or dataset, while preserving biological variance (age, sex, diagnosis, ...).
This is the prerequisite for any honest mega-analysis that pools individual-participant data (IPD) across sites or datasets.
This skill follows NeuroClaw hierarchy:
- Defines WHAT to do, not low-level implementation details.
- Does not execute direct shell commands itself.
- Delegates all execution via
claw-shell.
Research use only.
When to Use This Skill
Trigger this skill when the user asks for any of:
- "harmonize features across sites / scanners / datasets"
- "ComBat / ComBat-GAM / CovBat / neuroHarmonize / neuroCombat"
- "remove site effect / scanner effect / batch effect"
- "mega-analysis on ABIDE / ADHD-200 / ABCD / multi-site"
- "leave-site-out cross-validation"
- "site-stratified split"
- "cross-site generalization"
- "IPD pooling across cohorts"
Do NOT trigger for:
- Single-site, single-scanner studies (no batch variable exists)
- Pure preprocessing requests (delegate to
fmri-skill/smri-skill) - Model training itself (delegate to model skills via
run_models)
Position in the NeuroClaw Pipeline
[dataset-skill] → feature matrix + meta (site, scanner, age, sex, dx)
|
v
[harmonization-tool] ← this skill
|
v
harmonized feature matrix + same meta
|
v
[model-skill via run_models]
What ships with it
22 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- scripts/adapters/__init__.py 1.8 KB runs code
- scripts/adapters/base.py 1.2 KB runs code
- scripts/adapters/covbat_wrapper.py 2.7 KB runs code
- scripts/adapters/neurocombat_wrapper.py 2.9 KB runs code
- scripts/adapters/neuroharmonize_wrapper.py 4.1 KB runs code
- scripts/adapters/site_covar.py 3.2 KB runs code
- scripts/diagnostics.py 3.0 KB runs code
- scripts/extraction/atlas_registry.py 5.7 KB runs code
- scripts/extraction/extractors.py 4.0 KB runs code
- scripts/extraction/run_batch.py 7.7 KB runs code
- scripts/extraction/worker.py 10.0 KB runs code
- scripts/fetch_abide_rois_robust.py 4.9 KB runs code
- scripts/fetch_abide_rois.py 1.7 KB runs code
- scripts/harmonize.py 7.6 KB runs code
- scripts/io_schema.py 3.9 KB runs code
- scripts/loaders/__init__.py 214 B runs code
- scripts/loaders/abide_real.py 4.4 KB runs code
- scripts/loaders/adhd200_real.py 4.6 KB runs code
- scripts/pilot_abide_style.py 16 KB runs code
- scripts/splitters/__init__.py 264 B runs code
- scripts/splitters/leave_site_out.py 2.0 KB runs code
- scripts/splitters/site_stratified.py 2.6 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 275 lines · 119 tokens per session scan A 0b5800d4e097
harmonization-tool is a skill published in the GitHub repository CUHK-AIM-Group/NeuroClaw (83 stars, last pushed 2d ago), licensed MIT. It adds 119 tokens to every session and 3,869 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
arboreto
Infer gene regulatory networks (GRNs) from gene expression data using scalable algorithms (GRNBoost2, GENIE3). Use when analyzing transcriptomics data (bulk RNA-seq, single-cell RNA-seq) to identify transcription factor-target gene relationships and regulatory interactions. Supports distributed computation for…
torchdrug
Build and troubleshoot TorchDrug 0.2.1 workflows for molecular graphs, property prediction, self-supervised pretraining, molecule generation, retrosynthesis, protein representation learning, and knowledge graph reasoning. Use when code imports torchdrug or needs its datasets, models, tasks, or Engine.
deepspot-m
Generate transcriptome-wide virtual spatial transcriptomics from H&E histology with DeepSpot-M. Use when you need spatial gene expression in log1p-CPM for 224x224 tiles at about 20x, want to query protein-coding genes by symbol instead of a fixed panel, or want to run prediction across a whole slide after tiling with…
pyhealth
Build clinical/healthcare deep-learning pipelines with PyHealth — loading EHR/signal/imaging datasets (MIMIC-III/IV, eICU, OMOP, SleepEDF, ChestXray14, EHRShot), defining tasks (mortality, readmission, length-of-stay, drug recommendation, sleep staging, ICD coding, EEG events), instantiating models (Transformer…
pick-a-pii-model
Select an on-device OpenMed PII model from the committed registry by language, runtime format, and size budget, then require recall validation before deployment. Use when an agent must choose a local PII detector for CPU, Apple Silicon, or a mobile export without relying on live model discovery.
evo2
Score, embed, and generate DNA sequences with Evo 2, a long-context genomic foundation model. Use this skill when: (1) Computing per-nucleotide or per-sequence likelihoods for variant effect scoring, (2) Embedding genomic windows for downstream classification, (3) Generating DNA conditioned on a prefix, (4) Scoring…