Instructions file
Instructions for adrianco/retort, covering working in this repo, principle: one experiment at a time on this machine, publishing: blogs are one line per paragraph, experiment workflow and code layout — where cli commands live.
Instructions file
Instructions for adrianco/retort, covering working in this repo, principle: one experiment at a time on this machine, publishing: blogs are one line per paragraph, experiment workflow and code layout — where cli commands live.
Hook Claude Code
Runs before the context is compacted, running bd prime. From adrianco/retort.
Hook Claude Code
Runs when a session starts, running bd prime. From adrianco/retort.
Settings file Claude Code
Agent settings declaring 2 hook events (PreCompact, SessionStart).
Skill Claude CodeCodex
Use when working in a repository that uses bd or Beads for durable project task tracking, issue dependencies, blocker management, multi-session handoff, or shared work memory. Trigger when the user asks to find ready work, claim or close tasks, create follow-up work, inspect blockers, recover project context, or…
Skill Claude CodeCodex
Compare evaluated runs in a retort experiment along factor dimensions. Surfaces effects of each factor, aggregates across replicates, and highlights cells that diverge qualitatively — complementing (not replacing) retort's ANOVA analysis.
Skill Claude CodeCodex
Determine the TRUE cause of a failed retort run before attributing it. Ground-truth every failure (run its tests, read its agent logs, inspect its workspace) and classify it as an infrastructure false-fail, a genuine model miss, or an environment issue — never trust the gate verdict or a log signature alone. Use…
Skill Claude CodeCodex
Evaluate a single retort experiment run. Score the generated code against the task's TASK.md requirements, run its build and tests, compute metrics, and emit a structured evaluation report plus a machine-readable findings file.
Skill Claude CodeCodex
Aggregate a retort run's findings.jsonl into a machine-readable assessment.json summary with severity counts, penalty score, requirement coverage, and top findings.
Skill Claude CodeCodex
Summarize the architecture of code generated by a single retort run. Produces module-level structure, interfaces, and control flow in a form suitable for cross-run comparison — not a full codebase-summary.
Skill Claude CodeCodex
Refresh the data tables in optimal-blog.md from master.db. Checks the data for integrity problems FIRST, then runs the generator that picks per-language winners and splices every GEN-marked table, then reconciles the surrounding prose. Use after new experiment results land, or when the optimal-blog numbers are stale.