open science agents

3 tagged open science, measured the same way as everything else here.

analyzer

01

aipoch/open-science

Agent

After a valid blind comparison, unblind the result and explain why the winner performed better. Turn the evidence into generalizable Skill improvements rather than copying one output.

3.3k yesterday A 0 tokens original Apache-2.0

comparator

02

aipoch/open-science

Agent

Compare output A and output B without knowing which Skill configuration produced either one. Judge task completion and output quality, not presumed implementation quality.

3.3k yesterday A 0 tokens original Apache-2.0

grader

03

aipoch/open-science

Agent

Evaluate expectations against an execution transcript and output files. Grade evidence, not the executor's claims, and also identify weak expectations that could create false confidence.

3.3k yesterday A 0 tokens original Apache-2.0