analyzer
01Agent
After a valid blind comparison, unblind the result and explain why the winner performed better. Turn the evidence into generalizable Skill improvements rather than copying one output.
Open Science is an open-source, local-first, model-agnostic AI research workbench for reproducible scientific research on macOS, Windows, and Linux.
Agent
After a valid blind comparison, unblind the result and explain why the winner performed better. Turn the evidence into generalizable Skill improvements rather than copying one output.
Agent
Compare output A and output B without knowing which Skill configuration produced either one. Judge task completion and output quality, not presumed implementation quality.
Agent
Evaluate expectations against an execution transcript and output files. Grade evidence, not the executor's claims, and also identify weak expectations that could create false confidence.