evals

evals is a plugin for Claude Code from latestaiagents/agent-skills. Its manifest loads nothing; the 5 skills it bundles cost 482 tokens per session together, scanned A, original, MIT.

A guide to evaluating language-model output with test datasets, model-based reviewers, regression checks, and maintained reference examples.

In plain words
What is it for?
Use it to build evaluation datasets, run regression evaluations, maintain a trusted test set, and measure quality trade-offs.
Why use it?
It helps you detect when an AI system becomes less accurate and compare quality against cost before releasing changes.

Plugin for Claude Code

Written for Claude Code: a Claude Code plugin manifest.

Good fit Use it to build evaluation datasets, run regression evaluations, maintain a trusted test set, and measure quality trade-offs.

Compare 6 plugins from other repositories ↓
Install in Claude Code
/plugin marketplace add latestaiagents/agent-skills/plugin install evals
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Clone the repo
git clone --depth 1 https://github.com/latestaiagents/agent-skills

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for evals

README.md
[![agentmods](https://agentmods.dev/badge/plugins/latestaiagents/agent-skills/evals/github.svg)](https://agentmods.dev/plugins/latestaiagents/agent-skills/evals)
Your own site
<a href="https://agentmods.dev/plugins/latestaiagents/agent-skills/evals"><img src="https://agentmods.dev/badge/plugins/latestaiagents/agent-skills/evals/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for evals

Your own site · 80×15
<a href="https://agentmods.dev/plugins/latestaiagents/agent-skills/evals"><img src="https://agentmods.dev/badge/plugins/latestaiagents/agent-skills/evals.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session not measured What this adds to a session before it is invoked.
When invoked not measured The manifest loads nothing itself; its 5 skills cost 482 tokens a session between them.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Security

Grade A, and why

evals scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude-plugin/marketplace.json#evals · 21 lines

What it actually says

{
  "name": "evals",
  "source": "./skills/evals",
  "description": "Measure and improve LLM quality. LLM-as-judge, eval dataset design, regression evals, golden set maintenance, cost/quality tradeoff. 5 skills for serious eval engineering.",
  "version": "1.0.0",
  "author": {
    "name": "latestaiagents"
  },
  "homepage": "https://github.com/latestaiagents/agent-skills",
  "repository": "https://github.com/latestaiagents/agent-skills",
  "license": "MIT",
  "keywords": [
    "evals",
    "llm-judge",
    "regression-testing",
    "golden-set",
    "benchmark",
    "cost-quality"
  ],
  "category": "Quality Assurance"
}
Contents

What it installs

The manifest is a name and a version. 5 skills travel with it, and installing the plugin installs all of them — 482 tokens a session between them. Each is measured on its own page, and each can be installed alone.

Files

What ships with it

2 files beside marketplace.json#evals in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 21 lines scan A 4fdc91008d79

Subscribe to this mod's changes

evals is a plugin published in the GitHub repository latestaiagents/agent-skills (5 stars, last pushed 4mo ago), licensed MIT. Its token cost is not measured: this kind of file is read by the harness, not the model. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.