improvement skills

20 tagged improvement, measured the same way as everything else here.

Browse within: iteration 7artifact 6criteria 6evaluator 6generator 6prompt 5

trulens-diagnosis

01

truera/trulens

Skill Claude CodeCodex

Diagnose low evaluation scores and generate actionable improvement recommendations.

3.5k 4d ago A 17 tokens original MIT

simmer-judge-board

02

2389-research/simmer

Skill Claude CodeCodex

Judge board subskill for simmer. Dispatches a panel of judges with different lenses, runs one deliberation round where they challenge each other's scores, then synthesizes consensus scores + single ASI. Drop-in replacement for simmer-judge that produces identical output format. Do not invoke directly — dispatched by…

14 1mo ago A 76 tokens original MIT

simmer-setup

03

2389-research/simmer

Skill Claude CodeCodex

Setup subskill for simmer. Inspects the artifact or workspace, infers evaluation contracts and search space, proposes a complete assessment to the user, and produces a setup brief after confirmation. Conversational, not form-based — the agent does the work of understanding the problem, then presents what it found. Do…

14 1mo ago A 76 tokens original MIT

simmer

04

2389-research/simmer

Skill Claude CodeCodex

Use when user says "simmer this", "refine this", "hone this", "iterate on this", or asks to improve a specific artifact over multiple rounds. Runs an iterative refinement loop with investigation-first judges that read the code, understand the problem, and propose evidence-based improvements. Auto-selects single judge…

14 1mo ago A 124 tokens original MIT