smixs/skill-conductor

Architecture-first skill lifecycle for AI agents. BinEval binary scoring with threshold-blind, cross-family-calibrated judges, gated self-update loop, pressure testing, 10 authoring principles grounded in empirical research.

163Stars on the repository
7Mods indexed here, across every type
29d agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

analyzer

01

smixs/skill-conductor

Agent

Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.

163 29d ago A 0 tokens original MIT

bineval

02

smixs/skill-conductor

Agent

Evaluate a skill artifact with atomic binary yes/no questions, one answer (1/0) per question, each preceded by a written critique grounded in evidence from the skill's own files. Aggregate to per-dimension scores in [0,1]; the orchestrator turns your answers into the overall score and the pass/fail gate.

163 29d ago A 0 tokens original MIT

comparator

03

smixs/skill-conductor

Agent

Compare two outputs WITHOUT knowing which skill produced them, using binary yes/no questions.

163 29d ago A 0 tokens original MIT

grader

04

smixs/skill-conductor

Agent

Evaluate expectations against an execution transcript and outputs.

163 29d ago A 0 tokens original MIT