eval

A session-evaluation skill for scoring a completed agent session against a registered rubric, a fixed set of review criteria.

In plain words
What is it for?
Use it to run or review an evaluation, produce an evaluation report, or check whether a previous evaluation can be reproduced.
Why use it?
It provides a repeatable way to assess how an orchestrated work session was conducted and verify stored evaluation results.

Skill for Claude CodeCodexCursor

Part of the session-orchestrator plugin — 20 skills, 28 commands, 4 hooks shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/kanevry/session-orchestrator/eval
Any agent
npx skills add Kanevry/session-orchestrator --skill eval
Clone the repo
git clone --depth 1 https://github.com/Kanevry/session-orchestrator

Made for: Claude Code, Codex, Cursor.

Or install session-orchestrator, the plugin that ships this one along with the rest of its 20 skills, 28 commands, 4 hooks.

Per session 90 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 158 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00090 $0.00158
Opus 5 $0.00045 $0.00079
Sonnet 5 $0.00018 $0.00032
Haiku 4.5 $0.00009 $0.00016

Measured 3d ago against content hash 131ff70b3031, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.cursor/skills/eval/SKILL.md · 13 lines

What it actually says

eval

Canonical skill: skills/eval/SKILL.md

Read that file and follow it exactly. Resolve relative links against skills/eval/, not this wrapper.

Cursor has no Skill tool. Treat "invoke the eval skill" as: Read skills/eval/SKILL.md.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 3d ago First seen · 13 lines · 90 tokens per session scan A 131ff70b3031

Subscribe to this mod's changes

eval is a skill published in the GitHub repository Kanevry/session-orchestrator (49 stars, last pushed 5d ago), licensed MIT. It adds 90 tokens to every session and 158 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

semantix

Install and use the semantix memory kernel as a middleware in your agent: extract user preferences / workflows / experience from past sessions, retrieve and inject them on demand. One binary + your agent's own tools.

Gnosil/semantix · 47 tokens

semantix-guide

Troubleshoot and configure Semantix capabilities: Skills (project/custom/global/builtin priority, discovery dirs), Commands (override order, /dir:file naming), Hooks (11 events, automatic project loading, matchers, timeouts), MCP (semantix-agent.toml + .mcp.json + plugin packages, autostart), plugin packages…

Gnosil/semantix · 115 tokens

process-builder

Scaffold new babysitter process definitions following SDK patterns, proper structure, and best practices. Guides the 3-phase workflow from research to implementation.

a5c-ai/babysitter · 32 tokens

autoprompt

Explicit-only useful-first orchestration. Invoke /autoprompt to turn a mission into one executable roadmap, build dependency-safe lanes, and verify the result with independent reviewers. Never infer invocation from ordinary requests. Never resume from leftover artifacts without an explicit resume instruction.

Spielewoy/autoprompt-skill · 56 tokens

test-driven-development

Strict RED-GREEN-REFACTOR cycle enforcement. Tests are never skipped or deferred. Run mode only, never watch mode. Exit code evidence mandatory.

a5c-ai/babysitter · 35 tokens

product-brief-creation

Create comprehensive product briefs from market, domain, and technical research.

a5c-ai/babysitter · 0 tokens