Fill a slot-based question-framing schema (PICO, PECO, or SPIDER) from a paper's stated research question. Use this whenever the user wants a paper's research question structured into one of these standard clinical/qualitative-research question frames; this frames what question is being asked, it does not read or…
Select the minimal set of 1-3 verbatim sentences from a candidate paper/abstract sufficient to entail or refute an atomic claim (SciFact's rationale-selection step). Use this after claim-writing has produced an atomic claim, as the evidence-gathering step before claim-label-prediction; an empty rationale set is a…
Tactic: Grade an ML/CS paper's reproducibility configuration reporting as complete, partial, or none after checking that clinical appraisal tools do not apply. Use when the question is whether the work can be rerun.
Check whether a paper reports each item from PRISMA, CONSORT, STROBE, ARRIVE, SPIRIT, or TRIPOD (per whichever study-design-tool-gate dispatched to), citing where each item is or isn't addressed — including a/b sub-item hierarchy where the standard defines one. Use this after study-design-tool-gate has dispatched to…
(Proposal, unverified) Attempt to verify a paper's reported results by actually executing its released code/scripts against its own reported configuration — the only SOP in this package whose action type is code execution rather than text reading/judgment. Use this after unit-classification has extracted the paper's…
Judge a paper's stated research question against the FINER criteria (Feasible, Interesting, Novel, Ethical, Relevant) — five independent judgments with justification, evaluating the question itself, not the paper's results. Use this whenever the user wants to know whether a paper is asking a good research question…
(Proposal, unverified) Judge whether argumentative relations between unit-classification's rhetorical labels actually hold in a paper (e.g. is an AIM label adequately substantiated by BACKGROUND labels) — a second-order quality judgment over already-classified units, not raw text. Use this after unit-classification…
Keshav's second pass — a careful full read (ignoring proof/derivation detail) producing prose-level understanding sufficient to explain the paper's main contribution and evidence to a colleague. Use this after first-pass-skim, as the main content-grasping pass of the Keshav three-pass method; do not force its output…
Answer per-domain signalling questions (5-value scale: Yes/Probably yes/Probably no/No/No information) for RoB2, ROBINS-I, or QUADAS-2, per whichever variant study-design-tool-gate dispatched to. Use this after study-design-tool-gate has dispatched to one of these three tools; this SOP produces only the raw signalling…
Award NOS's (Newcastle-Ottawa Scale) stars item-by-item across Selection (up to 4), Comparability (up to 2), and Outcome/Exposure (up to 3) — a binary award-or-not action per item, distinct from a 5-value signalling judgment. Use this after study-design-tool-gate has dispatched to NOS, as the first step before…
Classify a paper's study design (RCT, cohort, case-control, diagnostic-accuracy, systematic-review, animal-study, prediction-model, etc., or notapplicable) and dispatch to the correct downstream bias-risk/quality/reporting tool and specific variant (CASP has 8 variants, JBI 6, RoB2 has parallel/cluster/crossover…
Sum NOS's item-level stars and bucket into good (≥7)/fair (4-6)/poor (≤3) — a fixed threshold lookup, structurally distinct from worst-case-lookup's take-the-worst-value approach. Use this after star-awarding has produced the per-item stars; this is NOS's terminal step.
Fill a paper's reported values into an already-given comparison-template attribute schema (e.g. Task/Dataset/Metric/Value) — the executable half of ORKG's comparison-template method. Use this when a template's attribute schema is already fixed and you need one paper's row filled in; this does NOT build new templates…
Keshav's third pass — the heaviest of the three, a full sentence-by-sentence re-read including proofs/derivations, attempting a virtual re-implementation of the paper to surface implicit assumptions and concrete improvement points. Use this after second-pass-grasp, as the terminal step of the Keshav three-pass method…
Classify each pre-segmented text unit independently against a fixed label set (Argumentative Zoning, CoreSC, PubMed-RCT, Swales move/step, CODA-19, TDMS, or CSFCube's facet labels), single-layer with no cross-unit dependency. Use this after unit-segmentation has split the text, whenever a sentence- or clause-level…
Split a paper's text into sentence- or clause-level units (with character offsets) for downstream classification, at a caller-specified granularity and scope (full text, abstract-only, or intro-only). Use this as the mandatory first step whenever any sentence/clause-level classification method (Argumentative Zoning…
Take the single most severe domain/item judgment as the overall verdict, for RoB2 (3-value), ROBINS-I (5-value), or AMSTAR-2 (pre-filtered by critical-domain status before worst-case). Use this after domain-level-judgment (for RoB2/ROBINS-I) or quality-appraisal-checklist (for AMSTAR-2) has produced per-domain/item…
Reconstruct a DARE skill repo's true use-dependency relations and render them as a self-contained, offline, Obsidian-style interactive HTML graph (pyvis / vis-network). Use this whenever the user wants to graph / map / visualize the skill dependencies of a repo or package, "画依赖图 / graph 化这个 repo / 把 skill 连边画出来 / 用…
Classify assumptions on 2 axes — load-bearing (how much conclusion depends on it) × vulnerable (how likely to be false). Focuses attention on High-Load × High-Vulnerable quadrant.