di37/EvalSurfer

Skill-first, agent-native evaluation protocol for AI apps

11Stars on the repository
1Mods indexed here, across every type
1mo agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

eval-surfer

01

di37/EvalSurfer

Skill Claude CodeCodex

Drive AI application evaluations using the EvalSurfer skill-first workflow. Use when creating AI eval rubrics, reviewing RAG outputs, checking agent tool use, assessing safety, or calculating operational metrics like latency, TTFT, inter-token latency, throughput (tokens per second), P99 tail latency, cost, cost per…

11 1mo ago A 81 tokens original MIT