Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/shihongdev/evalyn/evalyn-evalnpx skills add shihongDev/evalyn --skill evalyn-evalgit clone --depth 1 https://github.com/shihongDev/evalynWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00026 | $0.00792 |
| Opus 5 | $0.00013 | $0.00396 |
| Sonnet 5 | $0.00005 | $0.00158 |
| Haiku 4.5 | $0.00003 | $0.00079 |
Grade A, and why
evalyn-eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 110 lines — stays where its author put it; the contents beside it link to each section on GitHub.
evalyn-eval
Overview
Build a dataset from traces, auto-recommend metrics based on trace analysis, and run evaluation. This skill reads actual trace data to make metric recommendations rather than asking abstract questions.
Pre-flight
- Verify traces exist:
evalyn list-calls --limit 5
If no traces: "You need to instrument your agent first. Invoke evalyn-setup."
- Check if a dataset already exists:
ls data/*/dataset.jsonl 2>/dev/null
If dataset exists, skip to Step 2.
Step 1: Build Dataset
Identify the project name from the evalyn list-calls output (project column).
evalyn build-dataset --project <project-name>
Capture the output path - it prints "Wrote N items to ". Use this path for all subsequent commands.
Step 2: Auto-Recommend Metrics
Inspect a trace to understand the agent's behavior:
evalyn show-trace --last -v
Analyze the trace structure and recommend a bundle. Evalyn has 17 curated metric bundles:
| Trace Pattern | Recommended Bundle |
|---|---|
| Multiple tool calls, planning steps | orchestrator |
| Tool calls + multi-turn context | multi-step-agent |
| URLs or citations in output | research-agent |
| RAG retrieval spans, source docs | rag-qa |
| Conversational, multi-turn | chatbot |
| Code blocks in output | code-assistant |
| Short summary outputs | summarization |
| Educational/tutorial content | tutor |
| Content generation, blog posts | content-writer |
| Customer-facing Q&A | customer-support |
To see all available bundles:
evalyn suggest-metrics --mode bundle --help
Apply the recommended bundle:
evalyn suggest-metrics --dataset <path> --mode bundle --bundle <recommended>
Then expand coverage with LLM-based selection from the full 130+ metric registry:
evalyn suggest-metrics --dataset <path> --mode llm-registry --append
This two-pass approach gives a solid base (curated bundle) plus tailored additions (LLM picks from full registry).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 110 lines · 26 tokens per session scan A d2cc0731ea38
evalyn-eval is a skill published in the GitHub repository shihongDev/evalyn (257 stars, last pushed 2mo ago), licensed MIT. It adds 26 tokens to every session and 792 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
babysit-pr
Babysit a GitHub pull request after creation by continuously polling review comments, CI checks/workflow runs, and mergeability state until the PR is merged/closed or user help is required. Diagnose failures, retry likely flaky failures up to 3 times, auto-fix/push branch-related issues when appropriate, and keep…
imagegen
Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output…
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…
next-cache-components-optimizer
Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…