Compares several LLMs head to head on the same research goal using the co-scientist cross-model Elo bench, with optional gold-set scoring, and reads the results. Use when the user asks which model is best for hypothesis generation, wants to compare models or backends on a research task, asks to run a bench, or asks…
Improves an existing scientific hypothesis by combining it with another, simplifying it, making it feasible, or reimagining it out of the box, and records the result as a new hypothesis with its parents in the co-scientist database. Use when the user asks to improve, fix, merge, simplify, repair or rethink a…
Checks that the citations and evidence behind a hypothesis, review or write-up actually say what they are claimed to say, and reports the ones that do not. Use when the user asks to verify sources, check citations, confirm a claim is supported, or before publishing, sharing or acting on anything the co-scientist…
Compares two hypotheses head to head in a structured scientific debate and records the outcome as a tournament match, updating both Elo ratings in the co-scientist database. Use when the user asks which hypothesis is better, asks to compare, rank, or pit hypotheses against each other, or wants the tournament advanced…
Reads and explains the output of an AI co-scientist session, including the final research overview, the Elo-ranked hypotheses, and the reviews behind them. Use when the user asks what the co-scientist found, what the top hypotheses are, what a session produced, what an Elo score means here, or asks for the results to…
Reviews one scientific hypothesis for novelty, correctness and testability, or deep-verifies the assumptions it rests on, and records the review into the co-scientist research database so the tournament and meta-review count it. Use when the user asks to review, critique, check, stress-test or verify a hypothesis, or…
Starts, monitors, pauses, resumes and steers an AI co-scientist research session that generates and Elo-ranks novel scientific hypotheses from a research goal. Use when the user asks to run the co-scientist, start a research session, generate hypotheses for a scientific goal, check on a running session, or steer one…