analyzer
01Agent
After a valid blind comparison, unblind the result and explain why the winner performed better. Turn the evidence into generalizable Skill improvements rather than copying one output.
29 tagged ai for science, measured the same way as everything else here.
Browse within: llm-agents 12amd 11drug-discovery 11earth-science 11fine-tuning 11healthcare 11data-analysis 10instruments 10magnetism 10autonomous-research 5
Agent
After a valid blind comparison, unblind the result and explain why the winner performed better. Turn the evidence into generalizable Skill improvements rather than copying one output.
Agent
Compare output A and output B without knowing which Skill configuration produced either one. Judge task completion and output quality, not presumed implementation quality.
Agent
Evaluate expectations against an execution transcript and output files. Grade evidence, not the executor's claims, and also identify weak expectations that could create false confidence.
Agent Cursor
Score, diagnose, and gate a research report draft for grounded-review. Prefer a model different from the writer when available.
Agent Cursor
Apply reviewer-approved repairs to the research report draft for grounded-review while preserving substance.
Sibyl-Research-Team/AutoResearch-SibylSystem
Agent Claude Code
Sibyl Research System heavy-reasoning agent. Used for deep analysis tasks: synthesis, supervision, editing, critical review, and reflection.
Sibyl-Research-Team/AutoResearch-SibylSystem
Agent Claude Code
Sibyl Research System lightweight agent. Used for quick evaluation tasks: debate roles (optimist, skeptic, strategist), section critique, and cross-critique.
Sibyl-Research-Team/AutoResearch-SibylSystem
Agent Claude Code
Sibyl Research System standard agent. Used for literature research, planning, experiment design, and idea generation.
Agent
Delegate when drafting research communications, summaries, or reports for a non-specialist audience. Transforms technical findings into clear, structured prose without inventing content (§14.7).
Agent
Reads 3–7 key papers in depth and extracts claims, evidence, and methodology. Targets T1·T2 papers verified by citation-auditor (§14.7).
Agent
Broadly collects candidate papers using MCP connectors and skills. Generates 3–6 query families and assigns tier classifications (§14.7).
Agent
Submits a 2-node ORBIT-2 training (AMD Instinct MI355X) with PyTorch profiling and Omnistat user-mode telemetry, waits for completion, and writes manifest.json for downstream subagents.
Agent
Drive omnistat-inspect (PR #271) through the analyze-job phases on the user-mode VictoriaMetrics database, then map findings to the bottleneck taxonomy.
Agent
Independently re-derive the top 2-3 claims from omnistat/claims.json by issuing raw PromQL via curl against the same VictoriaMetrics endpoint, at the finest sampling step. Optionally probe one cheap remedy on a 1-node interactive srun.