benchmark-critic

benchmark-critic is a skill for Claude Code, Codex from cherninlab/ennodia. It costs 25 tokens per session (130 once invoked), scanned A, original, MIT.

A review of a software benchmark’s design or results for fairness, measurable scoring, information leakage, selective task choices, and repeatability.

In plain words
What is it for?
Use it to assess baselines, test fixtures, scoring rules, negative examples, raw outputs, and the smallest benchmark likely to produce useful evidence.
Why use it?
It helps identify when benchmark results may be misleading or impossible for others to check fairly.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to assess baselines, test fixtures, scoring rules, negative examples, raw…

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/cherninlab/ennodia/benchmark-critic
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add cherninlab/ennodia --skill benchmark-critic
Clone the repo
git clone --depth 1 https://github.com/cherninlab/ennodia

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for benchmark-critic

README.md
[![agentmods](https://agentmods.dev/badge/skills/cherninlab/ennodia/benchmark-critic.svg)](https://agentmods.dev/skills/cherninlab/ennodia/benchmark-critic)
Your own site
<a href="https://agentmods.dev/skills/cherninlab/ennodia/benchmark-critic"><img src="https://agentmods.dev/badge/skills/cherninlab/ennodia/benchmark-critic.svg" alt="Measured on agentmods" height="20"></a>
Per session 25 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 130 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00025 $0.00130
Opus 5 $0.00013 $0.00065
Sonnet 5 $0.00005 $0.00026
Haiku 4.5 $0.00003 $0.00013

Measured 7d ago against content hash d4b284769cd3, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-06, from the pricing page.

Security

Grade A, and why

benchmark-critic scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/benchmark-critic/SKILL.md · 21 lines

What it actually says

Benchmark Critic

Review benchmark ideas or results for credibility.

Look for:

  • a fair single-model baseline
  • a fair Ennodia or multi-agent condition
  • deterministic fixtures where possible
  • clear scoring rules before results are collected
  • leakage from oracle answers into prompts
  • cherry-picked tasks or missing negative examples
  • enough raw outputs to let others inspect the result

Recommend the smallest benchmark that can show useful signal without pretending to prove more than it does.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 7d ago First seen · 21 lines · 25 tokens per session scan A d4b284769cd3

Subscribe to this mod's changes

benchmark-critic is a skill published in the GitHub repository cherninlab/ennodia (5 stars, last pushed 2d ago), licensed MIT. It adds 25 tokens to every session and 130 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

agent-puzzles

Competitive puzzle arena for AI agents with timed solving, per-model leaderboards, and 5 categories (reverse captcha, geolocation, logic, science, code). Use when solving puzzles, tracking rankings, creating new challenges, or benchmarking agent capabilities.

ThinkOffApp/ide-agent-kit · 53 tokens

kotlin-testing

Kotlin testing patterns with Kotest, MockK, coroutine testing, property-based testing, and Kover coverage. Follows TDD methodology with idiomatic Kotlin practices.

affaan-m/ECC · 38 tokens

tdd-workflow

Use this skill when writing new features, fixing bugs, or refactoring code. Enforces test-driven development with 80%+ coverage including unit, integration, and E2E tests.

affaan-m/ECC · 43 tokens

react-testing

React component testing with React Testing Library, Vitest/Jest, MSW for network mocking, accessibility assertions with axe, and the decision boundary between component tests and Playwright/Cypress end-to-end runs. Use when writing or fixing tests for React components, hooks, or pages.

affaan-m/ECC · 59 tokens

python-testing

Python testing best practices using pytest including fixtures, parametrization, mocking, coverage analysis, async testing, and test organization. Use when writing or improving Python tests.

affaan-m/ECC · 35 tokens

react-patterns

React 18/19 patterns including hooks discipline, server/client component boundaries, Suspense + error boundaries, form actions, data fetching, state management decision trees, and accessibility-first composition. Use when writing or reviewing React components.

affaan-m/ECC · 49 tokens