Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Aperivue/medsci-skills --skill model-evaluationgit clone --depth 1 https://github.com/Aperivue/medsci-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/aperivue/medsci-skills/model-evaluation)<a href="https://agentmods.dev/skills/aperivue/medsci-skills/model-evaluation"><img src="https://agentmods.dev/badge/skills/aperivue/medsci-skills/model-evaluation/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/aperivue/medsci-skills/model-evaluation"><img src="https://agentmods.dev/badge/skills/aperivue/medsci-skills/model-evaluation.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00206 | $0.01802 |
| Opus 5 | $0.00103 | $0.00901 |
| Sonnet 5 | $0.00041 | $0.00360 |
| Haiku 4.5 | $0.00021 | $0.00180 |
Grade A, and why
model-evaluation scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 114 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Model-Evaluation Skill
Purpose
This skill makes a medical-imaging model's held-out evaluation task-correct and honest: the right metric for the task and the prevalence, with uncertainty, calibration, and subgroup performance. It emits a per-case metric table that the publication statistics build on, and gates the metric choice against Metrics Reloaded (Maier-Hein & Reinke et al., Nat Methods 2024) and CLAIM 2024.
It sits between /model-validation (which audits the split / design) and /analyze-stats (which owns
the comparative inference). It computes the imaging-specific per-case metrics (surface distances, FROC,
ECE of a softmax head); /analyze-stats owns DeLong / NRI / IDI / decision curves / MRMC. Like
/analyze-stats, it generates and executes code on your predictions — numbers are never hand-typed.
When to use
- You have held-out predictions + ground truth and need task-correct metrics with CIs, calibration, and subgroup slices, plus a per-case table for the manuscript statistics.
When NOT to use
- Auditing the validation design / leakage →
/model-validation. - DeLong / NRI / IDI / decision curves / MRMC reader study →
/analyze-stats. - Building / training the model →
/model-scaffold; LLM / MLLM →/mllm-eval. - Figure rendering →
/make-figures.
Workflow
Phase 1 — Fix the analysis unit and the task
State the task (segmentation / classification / detection / interactive / generative) and the analysis unit the metric must respect (per-patient vs per-lesion vs per-image). A per-lesion metric must not be reported as per-patient.
Phase 2 — Compute task-correct metrics
Generate evaluation code that computes, on the held-out predictions:
- segmentation: Dice/IoU and a boundary metric (HD95 / NSD), per structure not only a global mean, with bootstrap 95% CIs.
- classification: AUROC and AUPRC with bootstrap CIs, sensitivity/specificity, and PPV/NPV at the deployment prevalence (not a balanced set).
- detection: FROC / mAP with the IoU match criterion stated.
- interactive / promptable segmentation (SAM2 / MedSAM2 / nnInteractive): the segmentation
metrics above plus the interaction axis — Dice-vs-interactions / number-of-clicks (NoC) to a
target threshold, initial-vs-converged (or peak) Dice, and per-case interaction/inference time
(see the metric guide; the study design is in
/design-study+/model-validation). - generative / synthesis (image generation or modification): full-reference similarity
(MSE/RMSE/PSNR/SSIM) or no-reference quality (SNR/CNR, standardized visual scores), plus a
downstream-task evaluation — image quality is not clinical utility (Park et al., Radiol Med 2024).
For multiclass classification, state the aggregation scheme (one-vs-rest / macro / micro /
pairwise / Obuchowski); time-to-event discrimination (Harrell's C, time-dependent ROC) is handed
to
/analyze-stats. Add calibration (reliability diagram / ECE) and subgroup slices (the Model Card Factors). See${CLAUDE_SKILL_DIR}/references/metric_guide.md. Emit a per-case CSV for/analyze-stats.
What ships with it
18 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- references/metric_guide.md 6.0 KB
- references/metric_selection_grounding.md 10 KB
- scripts/check_metric_reporting.py 18 KB runs code
- scripts/metric_reporting_challenge/fixture/clf_bad.md 84 B
- scripts/metric_reporting_challenge/fixture/clf_good.md 186 B
- scripts/metric_reporting_challenge/fixture/det_good_wrapped.md 285 B
- scripts/metric_reporting_challenge/fixture/det_no_iou.md 224 B
- scripts/metric_reporting_challenge/fixture/generative_bad.md 134 B
- scripts/metric_reporting_challenge/fixture/generative_good.md 437 B
- scripts/metric_reporting_challenge/fixture/interactive_bad.md 128 B
- scripts/metric_reporting_challenge/fixture/interactive_good.md 500 B
- scripts/metric_reporting_challenge/fixture/multiclass_bad.md 128 B
- scripts/metric_reporting_challenge/fixture/seg_bad.md 119 B
- scripts/metric_reporting_challenge/fixture/seg_good.md 207 B
- scripts/metric_reporting_challenge/problem.md 1.6 KB
- scripts/metric_reporting_challenge/verify.sh 2.6 KB runs code
- skill.yml 2.9 KB
- tests/test_metric_reporting.sh 6.7 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 114 lines · 206 tokens per session scan A b39293e54402
model-evaluation is a skill published in the GitHub repository Aperivue/medsci-skills (290 stars, last pushed yesterday), licensed MIT. It adds 206 tokens to every session and 1,802 once invoked, about $0.0010 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
clinical-trials-search
Search ClinicalTrials.gov with natural language queries. Find clinical trials, enrollment, and outcomes using Valyu semantic search.
shidi
A Chinese-language research assistant role that turns a user’s ideas into literature reviews, experiment plans, figures and organised data.
clinical-research
Use when designing a prospective clinical study before submission — selecting and classifying endpoints (primary / key-secondary / exploratory, with surrogate-endpoint flagging), estimating sample size and power for two-arm designs (means / proportions / survival), or scoring a study plan for feasibility and a GO /…
sr-search-record
An automated literature-review search and screening workflow using OpenAlex, an open database and API for scholarly research, and Zotero, a reference manager.
critical-paper-reading
A paper-reading workflow for critically examining one research paper from a PDF or Zotero item. Zotero is a tool for organising research papers and references.
knowledge-module-gen
A workflow for generating knowledge-module files from a list of academic reading modules. It reads full papers from Zotero, NotebookLM, or local PDFs and records page references and execution checks.