model-calibration-curve

model-calibration-curve is a skill for Claude Code, Codex from aipoch/medical-research-skills. It costs 64 tokens per session (2,633 once invoked), scanned A, original, MIT.

A statistical check for a survival model, which predicts how likely patients are to remain alive or event-free at set times. It compares those predictions with observed patient outcomes using bootstrap calibration curves, which estimate how reliable the comparison is.

In plain words
What is it for?
Use it to assess one-, two-, or three-year survival predictions from a clinical CSV file and export calibration results with a PDF chart.
Why use it?
It shows whether predicted probabilities match what happened in the clinical data, rather than only showing whether patients were ranked correctly.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Needs its repository: it runs a file that does not travel with it, so clone the repository first. The line is --output_dir ./output/.

Good fit Use it to assess one-, two-, or three-year survival predictions from a clinical CSV file and export calibration results with a PDF chart.

Compare 6 skills from other repositories ↓
About the project

Medical Research Agent Skills is a library of agent instructions for medical and biomedical research, covering evidence analysis, study protocol design, data analysis, and academic writing. Researchers use it to guide compatible coding agents through common scientific workflows. The catalogue contains many of the library's skills and commands.

aipoch/medical-research-skills · 1,860 stars · on GitHub · aipoch.com

Install

Getting it into your agent

It runs from inside its repository, so the clone comes first — what it calls does not travel with the file alone.

Clone the repo
git clone --depth 1 https://github.com/aipoch/medical-research-skills
agentmods
npx agentmods add skills/aipoch/medical-research-skills/model-calibration-curve

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for model-calibration-curve

README.md
[![agentmods](https://agentmods.dev/badge/skills/aipoch/medical-research-skills/model-calibration-curve/github.svg)](https://agentmods.dev/skills/aipoch/medical-research-skills/model-calibration-curve)
Your own site
<a href="https://agentmods.dev/skills/aipoch/medical-research-skills/model-calibration-curve"><img src="https://agentmods.dev/badge/skills/aipoch/medical-research-skills/model-calibration-curve/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for model-calibration-curve

Your own site · 80×15
<a href="https://agentmods.dev/skills/aipoch/medical-research-skills/model-calibration-curve"><img src="https://agentmods.dev/badge/skills/aipoch/medical-research-skills/model-calibration-curve.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 64 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,633 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • NVIDIA SkillSpector warn 7 Sept 2026
SkillSpector: 1 finding, up to high

These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →

  • high Data Exfiltration · line 173
    Code or instructions that leak agent conversation context to external services, potentially exposing sensitive user interactions.
    Fix: Remove any code that sends prompts, responses, or session data externally. Preserve user privacy; never exfiltrate conversation content.
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00064 $0.02633
Opus 5 $0.00032 $0.01316
Sonnet 5 $0.00013 $0.00527
Haiku 4.5 $0.00006 $0.00263

Measured 13d ago against content hash 1dbdf02a16d4, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-12, from the pricing page.

Security

Grade A, and why

model-calibration-curve scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 13d ago.

The scan reads SKILL.md. This mod also ships 1 executable file (tests/run_smoke_test.sh), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

awesome-med-research-skills/Data Analysis/model-calibration-curve/SKILL.md · 306 lines

How it starts

The opening of the file, as written. The whole thing — 306 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Model Calibration Curve

When to Use

Use this skill when you need to:

  • validate a survival model with bootstrap calibration curves;
  • compare predicted and observed survival probabilities at multiple horizons;
  • export calibration statistics together with a PDF visualization.

Typical user requests:

  • "Generate 1-, 2-, and 3-year calibration curves for this prognosis model."
  • "Check whether the Cox model built from age, gender, and risk is well calibrated."
  • "Export calibration statistics and a calibration PDF from this clinical cohort."

When Not to Use

Do not use this skill for:

  • nomogram construction;
  • univariate or multivariable Cox feature screening;
  • ROC, calibration-free discrimination, or decision-curve analysis;
  • non-survival endpoints or multiclass classification tasks.

When to Read External Files

Situation File to Read Purpose
Need algorithm details references/algorithm.md Statistical method and formulas
Need to run analysis scripts/main.R Get the complete command
Encounter errors references/troubleshooting.md Find solutions
Need CLI examples references/cli-guide.md Parameter usage examples

Input Validation

This skill accepts:

  • A clinical CSV file with sample IDs as row names, survival time, event indicator, and pre-selected prognostic features
  • Requests to assess calibration of a survival (Cox) model via bootstrap resampling at one or more prediction horizons

If the user's request does not involve survival model calibration from a clinical CSV file — for example, asking to construct a nomogram, screen Cox features, generate an ROC curve, analyze a decision curve, or work with non-survival outcomes — do not proceed with this workflow. Instead respond:

"model-calibration-curve is designed to validate survival model calibration by generating bootstrap calibration curves from a clinical CSV file. Your request appears to be outside this scope. Please use a nomogram-construction skill for nomogram building, a roc-diagnostic-performance skill for ROC analysis, or a decision-curve-analysis skill for DCA."

Read the full file on GitHub · 306 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 13d ago First seen · 306 lines · 64 tokens per session scan A 1dbdf02a16d4

Subscribe to this mod's changes

model-calibration-curve is a skill published in the GitHub repository aipoch/medical-research-skills (1,860 stars, last pushed 1mo ago), licensed MIT. It adds 64 tokens to every session and 2,633 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories