inspect-evaluation-harness-starter

inspect-evaluation-harness-starter is a skill for Claude Code, Codex from ma-compbio-lab/SkillFoundry. It costs 0 tokens per session (281 once invoked), scanned A, original, Apache-2.0.

A local evaluation tool for testing scientific agents on a fixed set of small example tasks. It compares two deterministic solvers and saves accuracy summaries and Inspect log files.

In plain words
What is it for?
Use it to run toy evaluations, compare a candidate solver with a weaker baseline, and inspect the recorded results.
Why use it?
It provides a repeatable way to compare agent behavior without needing external model credentials.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Needs its repository: it runs a file that does not travel with it, so clone the repository first. The line is ./slurm/envs/agents/bin/python skills/scientific-agents-and-automation/inspect-evaluation-harness-starter/scripts/run_inspect_evaluation_harness.py \.

Good fit Use it to run toy evaluations, compare a candidate solver with a weaker baseline, and inspect the recorded results.

Compare 6 skills from other repositories ↓
Install

Getting it into your agent

It runs from inside its repository, so the clone comes first — what it calls does not travel with the file alone.

Clone the repo
git clone --depth 1 https://github.com/ma-compbio-lab/SkillFoundry
agentmods
npx agentmods add skills/ma-compbio-lab/skillfoundry/inspect-evaluation-harness-starter

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for inspect-evaluation-harness-starter

README.md
[![agentmods](https://agentmods.dev/badge/skills/ma-compbio-lab/skillfoundry/inspect-evaluation-harness-starter/github.svg)](https://agentmods.dev/skills/ma-compbio-lab/skillfoundry/inspect-evaluation-harness-starter)
Your own site
<a href="https://agentmods.dev/skills/ma-compbio-lab/skillfoundry/inspect-evaluation-harness-starter"><img src="https://agentmods.dev/badge/skills/ma-compbio-lab/skillfoundry/inspect-evaluation-harness-starter/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for inspect-evaluation-harness-starter

Your own site · 80×15
<a href="https://agentmods.dev/skills/ma-compbio-lab/skillfoundry/inspect-evaluation-harness-starter"><img src="https://agentmods.dev/badge/skills/ma-compbio-lab/skillfoundry/inspect-evaluation-harness-starter.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 0 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 281 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00000 $0.00281
Opus 5 $0.00000 $0.00140
Sonnet 5 $0.00000 $0.00056
Haiku 4.5 $0.00000 $0.00028

Measured 6d ago against content hash 42c3a82c2d23, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-09, from the pricing page.

Security

Grade A, and why

inspect-evaluation-harness-starter scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

The scan reads SKILL.md. This mod also ships 2 executable files (scripts/run_inspect_evaluation_harness.py, tests/test_run_inspect_evaluation_harness.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/scientific-agents-and-automation/inspect-evaluation-harness-starter/SKILL.md · 30 lines

What it actually says

Inspect Evaluation Harness Starter

Use this skill to run a deterministic Inspect evaluation harness over toy scientific-agent cases and compare a candidate solver against a weaker baseline.

What This Skill Does

  • defines a small local Inspect task set without external model credentials
  • evaluates two deterministic solver variants on the same cases
  • writes machine-readable accuracy summaries plus Inspect log files for both runs

When To Use It

  • when you need a runnable evaluation-harnesses-for-scientific-agents starter
  • when you want a local Inspect example before wiring in real agents or model-backed solvers
  • when you need a stable comparison harness for repository tests

Run

./slurm/envs/agents/bin/python skills/scientific-agents-and-automation/inspect-evaluation-harness-starter/scripts/run_inspect_evaluation_harness.py \
  --cases skills/scientific-agents-and-automation/inspect-evaluation-harness-starter/examples/toy_eval_cases.json \
  --summary-out scratch/agents/inspect_evaluation_harness_summary.json \
  --log-dir scratch/agents/inspect-eval-logs

Notes

  • This starter intentionally avoids external model APIs so it can run in the repository sandbox.
  • The candidate and baseline solvers are both deterministic; the purpose is to verify the harness and comparison surface, not to benchmark large models.
Files

What ships with it

6 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 30 lines · 0 tokens per session scan A 42c3a82c2d23

Subscribe to this mod's changes

inspect-evaluation-harness-starter is a skill published in the GitHub repository ma-compbio-lab/SkillFoundry (38 stars, last pushed 4mo ago), licensed Apache-2.0. It costs nothing until one of its globs matches a file; then it loads 281 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.

Related

Other skills, from other repositories

arxiv-search

Skill "arxiv-search" from TaewoooPark/MagLab, covering category scoping, query families (3–6 per topic), version and doi resolution, tier classification and output contract.

TaewoooPark/MagLab · 68 tokens

revision-letter

Point-by-point peer-review response letter workflow for journal resubmission. Invokes RevisionLetterAgent to quote each reviewer comment verbatim, draft a response, and add a change-location marker. Outputs carry HUMAN REVIEW REQUIRED; no auto-send. A DOI or manuscript location is required for every factual response…

TaewoooPark/MagLab · 68 tokens

literature-search

Broad literature search for magnetism & spintronics — OpenAlex REST query strategy, query family generation, tier classification, and evidencematrix construction (§14.3·§14.7). Activated by the maglab lit search command and the research orchestration search-scout agent.

TaewoooPark/MagLab · 60 tokens

physics-oracle

Use when validating the dimensional, range, and conservation-law plausibility of magnetic physics quantities, or when performing deterministic physics formula calculations and unit conversions. Gilbert damping 0≤α≤1 check, M≤Ms, exchange length and domain wall width calculations, Oe↔A/m↔T·emu/cm³↔A/m·Jij meV↔K…

TaewoooPark/MagLab · 84 tokens

statistical-experimental-evaluation

Design and run statistical experiments that test the formal problem, proposed methods, theoretical predictions, baselines, and ablations.

aiming-lab/AutoResearchClaw · 31 tokens

batch-effect-correction

Use when correcting batch effects in merged bulk expression matrices with sample-level batch metadata while preserving biological group structure and generating before-and-after QC plots. NOT for: single-cell integration, raw FASTQ processing, differential expression without batch labels, or datasets without…

aipoch/medical-research-skills · 57 tokens