shihongDev/evalyn

Evalyn — a lightweight, local-first eval framework for your GenAI app

257Stars on the repository
6Mods indexed here, across every type
2mo agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

evalyn-analyze

01

shihongDev/evalyn

Skill Claude CodeCodex

Use when analyzing evalyn evaluation results, investigating failures, comparing runs, or understanding agent performance.

257 2mo ago A 23 tokens original MIT

evalyn-calibrate

02

shihongDev/evalyn

Skill Claude CodeCodex

Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed.

257 2mo ago A 28 tokens original MIT

evalyn-eval

03

shihongDev/evalyn

Skill Claude CodeCodex

Use when building evaluation datasets, selecting metrics, or running evaluations on an LLM agent project with evalyn.

257 2mo ago A 26 tokens original MIT

evalyn-setup

04

shihongDev/evalyn

Skill Claude CodeCodex

Use when setting up evalyn evaluation for an LLM agent project, instrumenting agent code, or adding the evalyn decorator.

257 2mo ago A 30 tokens original MIT

evalyn

05

shihongDev/evalyn

Skill Claude CodeCodex

Use to evaluate an LLM agent with evalyn. Orchestrates the full pipeline: install, instrument, trace, build dataset, suggest metrics, run eval, analyze, calibrate.

257 2mo ago A 41 tokens original MIT