homemade-software-inc/completion-kit

Your prompts need tests too. Run prompts against real datasets, score outputs with LLM judges, version everything, and compare runs to see what got better.

3Stars on the repository
3Mods indexed here, across every type
21d agoLast push, which is what freshness is scored on
noneNo LICENSE: all rights reserved, so bodies are not copied

evals

01

homemade-software-inc/completion-kit

MCP server Claude CodeCodexCursor +2

Prompt evals over MCP: run a prompt on your dataset, score each output 1-5 with an LLM judge. Remote server at completionkit.com.

3 21d ago A tokens not measured