shihongDev

6 mods across 1 repository, 257 stars between them.

evalyn CLAUDE.md

01

shihongDev/evalyn

Instructions file

Instructions for shihongDev/evalyn, a project described as: Evalyn — a lightweight, local-first eval framework for your GenAI app.

257 2mo ago B 1,345 tokens original MIT

evalyn-analyze

02

shihongDev/evalyn

Skill Claude CodeCodex

Use when analyzing evalyn evaluation results, investigating failures, comparing runs, or understanding agent performance.

257 2mo ago A 23 tokens original MIT

evalyn-calibrate

03

shihongDev/evalyn

Skill Claude CodeCodex

Use when LLM judges need calibration, evaluation metrics seem misaligned with expectations, or annotation and judge tuning is needed.

257 2mo ago A 28 tokens original MIT

evalyn-eval

04

shihongDev/evalyn

Skill Claude CodeCodex

Use when building evaluation datasets, selecting metrics, or running evaluations on an LLM agent project with evalyn.

257 2mo ago A 26 tokens original MIT

evalyn-setup

05

shihongDev/evalyn

Skill Claude CodeCodex

Use when setting up evalyn evaluation for an LLM agent project, instrumenting agent code, or adding the evalyn decorator.

257 2mo ago A 30 tokens original MIT

evalyn

06

shihongDev/evalyn

Skill Claude CodeCodex

Use to evaluate an LLM agent with evalyn. Orchestrates the full pipeline: install, instrument, trace, build dataset, suggest metrics, run eval, analyze, calibrate.

257 2mo ago A 41 tokens original MIT