analyzer
01Agent
Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.
Architecture-first skill lifecycle for AI agents. BinEval binary scoring with threshold-blind, cross-family-calibrated judges, gated self-update loop, pressure testing, 10 authoring principles grounded in empirical research.
Agent
Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.
Agent
Evaluate a skill artifact with atomic binary yes/no questions, one answer (1/0) per question, each preceded by a written critique grounded in evidence from the skill's own files. Aggregate to per-dimension scores in [0,1]; the orchestrator turns your answers into the overall score and the pass/fail gate.
Agent
Compare two outputs WITHOUT knowing which skill produced them, using binary yes/no questions.
Agent
Evaluate expectations against an execution transcript and outputs.