create-eval

create-eval is a skill for Claude Code, Codex from UKGovernmentBEIS/inspect_evals. It costs 46 tokens per session (284 once invoked), scanned A, original, MIT.

A redirect for creating new Inspect Evals evaluations, which are tests used to measure how well AI systems perform specific tasks.

In plain words
What is it for?
It points evaluation authors to the inspect-evals-template and explains how the old repository handles registration instead.
Why use it?
It prevents work from being started in the old repository, where new evaluations are no longer created.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/ukgovernmentbeis/inspect_evals/create-eval
Any agent
npx skills add UKGovernmentBEIS/inspect_evals --skill create-eval
Clone the repo
git clone --depth 1 https://github.com/UKGovernmentBEIS/inspect_evals

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for create-eval

README.md
[![agentmods](https://agentmods.dev/badge/skills/ukgovernmentbeis/inspect_evals/create-eval.svg)](https://agentmods.dev/skills/ukgovernmentbeis/inspect_evals/create-eval)
Your own site
<a href="https://agentmods.dev/skills/ukgovernmentbeis/inspect_evals/create-eval"><img src="https://agentmods.dev/badge/skills/ukgovernmentbeis/inspect_evals/create-eval.svg" alt="Measured on agentmods" height="20"></a>
Per session 46 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 284 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00046 $0.00284
Opus 5 $0.00023 $0.00142
Sonnet 5 $0.00009 $0.00057
Haiku 4.5 $0.00005 $0.00028

Measured 5d ago against content hash ae770931ad6c, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

create-eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/skills/create-eval/SKILL.md · 20 lines

What it actually says

Create Evaluation

New evaluations are no longer created in the inspect_evals repository. Since May 2026, evaluations live in their own standalone repositories. This repo now serves as a register that points to upstream eval repos.

This skill has been moved to inspect-evals-template: https://github.com/Generality-Labs/inspect-evals-template/tree/main/.claude/skills/create-eval. If a user triggers this skill, inform them to go to the template to run or access the skill there.

If a user asks to create an eval, inform them that Inspect Evals no longer accepts code submissions for new evals and suggest the inspect-evals-template for guidance on creating evals.

Quick reference

Step Where
Create a new eval inspect-evals-template
Register a completed eval This repo — use the Prepare Evaluation For Submission skill
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 20 lines · 46 tokens per session scan A ae770931ad6c

Subscribe to this mod's changes

create-eval is a skill published in the GitHub repository UKGovernmentBEIS/inspect_evals (658 stars, last pushed yesterday), licensed MIT. It adds 46 tokens to every session and 284 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

systematic-debugging

Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.

obra/superpowers · 21 tokens

local-ai-agents

Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…

microsoft/ai-agents-for-beginners · 200 tokens

next-cache-components-adoption

Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…

vercel/next.js · 95 tokens

chat-pet-sprite-creation

Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.

microsoft/vscode · 53 tokens

cpu-profile-analysis

Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…

microsoft/vscode · 71 tokens

insight-error-page

Write or audit an insight-kind error page for the Next.js dev overlay. Use when creating a new errors/ .mdx page, auditing an existing one, or checking that a page matches the framework fix cards. Covers page structure, title alignment, FixCard cards with Copy prompt button, code snippets, terminology verification…

vercel/next.js · 83 tokens