JarvixGaby/eval-skill

Eval-skill tests any skill against itself — no scenarios to write, just a blind side-by-side report showing whether it's actually worth using

2Stars on the repository
4Mods indexed here, across every type
1mo agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

analyzer

01

JarvixGaby/eval-skill

Agent

Explain why blinded evaluation results occurred after identities are revealed. This stage may read the target skills, labelkey.json, sanitized and raw outputs, transcripts, grading files, metrics, timing, and comparison.json.

2 1mo ago A 0 tokens original MIT

comparator

02

JarvixGaby/eval-skill

Agent

Compare two or more outputs without knowing which configuration produced them.

2 1mo ago A 0 tokens original MIT

grader

03

JarvixGaby/eval-skill

Agent

Evaluate expectations against an execution transcript and outputs.

2 1mo ago A 0 tokens copy · 97% MIT