Skill Claude CodeCodex
Diagnose low evaluation scores and generate actionable improvement recommendations.
20 tagged improvement, measured the same way as everything else here.
Browse within: iteration 7artifact 6criteria 6evaluator 6generator 6prompt 5
Skill Claude CodeCodex
Diagnose low evaluation scores and generate actionable improvement recommendations.
Skill Claude CodeCodex
Judge board subskill for simmer. Dispatches a panel of judges with different lenses, runs one deliberation round where they challenge each other's scores, then synthesizes consensus scores + single ASI. Drop-in replacement for simmer-judge that produces identical output format. Do not invoke directly — dispatched by…
Skill Claude CodeCodex
Setup subskill for simmer. Inspects the artifact or workspace, infers evaluation contracts and search space, proposes a complete assessment to the user, and produces a setup brief after confirmation. Conversational, not form-based — the agent does the work of understanding the problem, then presents what it found. Do…
Skill Claude CodeCodex
Use when user says "simmer this", "refine this", "hone this", "iterate on this", or asks to improve a specific artifact over multiple rounds. Runs an iterative refinement loop with investigation-first judges that read the code, understand the problem, and propose evidence-based improvements. Auto-selects single judge…