code optimization agents

3 tagged code optimization, measured the same way as everything else here.

benchmark-reviewer

01

evo-hq/evo

Agent

Reviews an evo benchmark in two modes. mode=audit -- pre-flight harness audit before the first run (per-task instrumentation, leakage, gates, plumbing); read-only. mode=review-experiment -- post-commit per-task failure analysis for a specific experiment; reads per-task traces and the eval-runner log, writes per-task…

1.4k +6 1mo ago A 112 tokens original Apache-2.0

ideator

02

evo-hq/evo

Agent

Generates ranked experiment proposals for the evo orchestrator. Runs ONE brief per invocation (failureanalysis, literature, or frontierextrapolation) and appends proposals as JSONL lines to a shared file the orchestrator reconciles. Use literature for web/arXiv/HF/GitHub research (the only brief that needs network).…

1.4k +6 1mo ago A 142 tokens original Apache-2.0

verifier

03

evo-hq/evo

Agent

Read-only audit of one evo experiment for design-time cheating (pre-phase) or result-time validity (post-phase). Catches test-set leakage in training data, subsetted eval commands, missing gates for new artifacts, generic hypotheses, cache short-circuits, fake artifacts, and score-reproducibility failures. Returns…

1.4k +6 1mo ago A 146 tokens original Apache-2.0