Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add leonardodalinky/SciDER --skill literature-review-agentgit clone --depth 1 https://github.com/leonardodalinky/SciDERWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/leonardodalinky/scider/literature-review-agent)<a href="https://agentmods.dev/skills/leonardodalinky/scider/literature-review-agent"><img src="https://agentmods.dev/badge/skills/leonardodalinky/scider/literature-review-agent/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/leonardodalinky/scider/literature-review-agent"><img src="https://agentmods.dev/badge/skills/leonardodalinky/scider/literature-review-agent.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00132 | $0.03852 |
| Opus 5 | $0.00066 | $0.01926 |
| Sonnet 5 | $0.00026 | $0.00770 |
| Haiku 4.5 | $0.00013 | $0.00385 |
Grade A, and why
literature-review-agent scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
This is a copy
98% identical to literature-review-agent — 1 line differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
How it starts
The opening of the file, as written. The whole thing — 358 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Literature Review Agent (Step 3)
Faithful implementation of the Hybrid Literature Agent from PaperOrchestra (Song et al., 2026, arXiv:2604.05018, §4 Step 3, App. D.3, App. F.1 p.46).
Cost: ~20–30 LLM calls. This is one of the two longest steps (the other is plotting). Wall-time floor is set by Semantic Scholar's 1 QPS verification limit.
Inputs
workspace/outline.json— specificallyintro_related_work_planwith the Introduction search directions and the 2-4 Related Work methodology clustersworkspace/inputs/conference_guidelines.md— used to derivecutoff_dateworkspace/inputs/idea.md,workspace/inputs/experimental_log.md— for framing the Intro and grounding the Related Work positioning
Outputs
workspace/citation_pool.json— verified Semantic Scholar metadata for every paper that survived verificationworkspace/refs.bib— BibTeX file generated from the verified poolworkspace/drafts/intro_relwork.tex— drafted Introduction and Related Work sections, written into the template, with the rest of the template preserved verbatim
Two-phase pipeline (App. D.3)
PHASE 1 — Parallel Candidate Discovery
For each search direction in introduction_strategy.search_directions:
For each limitation_search_query in each related_work cluster:
- Use the host's web search tool to discover up to ~10 candidate papers.
- Run up to 10 discovery queries in parallel (host-permitting).
- Collect (title, snippet, url) tuples — no verification yet.
→ PRE-DEDUP before Phase 2 (see Step 1.5 below)
PHASE 2 — Sequential Citation Verification (1 QPS, with cache)
For each candidate (after pre-dedup), sequentially:
0. Check s2_cache.json first (scripts/s2_cache.py --check).
If HIT: use cached response, skip live S2 call. No throttle needed.
If MISS: proceed with live request below.
1. Query Semantic Scholar by title:
GET https://api.semanticscholar.org/graph/v1/paper/search?query=<title>
&fields=title,abstract,year,authors,venue,externalIds&limit=5
(Public endpoint, no key. Throttle to 1 QPS for live requests only.)
2. Store the S2 response in cache: s2_cache.py --store.
3. Pick the top hit. Check Levenshtein title ratio against the original
candidate title. If ratio < 70: discard.
4. Bonus: if year and venue exactly align with hints, add a +5 point
match-quality bonus.
5. Require: abstract is non-empty.
6. Require: paper.year (or month if known) strictly predates cutoff_date.
Months default to day-1: e.g., "October 2024" → 2024-10-01.
7. If all checks pass, add to verified pool.
After all candidates are verified, dedup by Semantic Scholar paperId.
What ships with it
17 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- references/citation-density-rule.md 2.8 KB
- references/discovery-pipeline.md 4.8 KB
- references/exa-search-cookbook.md 9.1 KB
- references/prompt.md 3.2 KB
- references/s2-api-cookbook.md 4.1 KB
- references/verification-rules.md 4.2 KB
- scripts/bibtex_format.py 5.4 KB runs code
- scripts/check_cutoff.py 2.2 KB runs code
- scripts/citation_coverage.py 3.1 KB runs code
- scripts/dedupe_by_id.py 3.1 KB runs code
- scripts/exa_search.py 5.8 KB runs code
- scripts/levenshtein_match.py 2.1 KB runs code
- scripts/pre_dedup_candidates.py 4.9 KB runs code
- scripts/s2_cache.py 3.5 KB runs code
- scripts/s2_search.py 7.1 KB runs code
- scripts/sync_keys.py 3.9 KB runs code
- scripts/validate_pool.py 4.8 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 358 lines · 132 tokens per session scan A b0fde92518d4
literature-review-agent is a skill published in the GitHub repository leonardodalinky/SciDER (88 stars, last pushed 3mo ago), licensed Apache-2.0. It adds 132 tokens to every session and 3,852 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. It is 98% identical to literature-review-agent, differing in 1 line, and is treated as a copy.
Other skills, from other repositories
drawio-reconstruction
Reconstructs reference images into high-fidelity, editable Draw.io files with rendered previews: native Draw.io elements carry text and structure, SVG covers simple icons that match the reference, and cropped or transparent PNGs preserve complex visuals. Use when the user wants a diagram image, research figure…
idea-evaluator
Evaluates a preliminary research idea against a five-dimension framework (Higher, Faster, Stronger, Cheaper, Broader) plus idea-lifecycle and student-capability matching, paradigm-shift probing, and a fatal-flaws audit. Returns a reviewer-style verdict; non-STEM ideas route to substitute frameworks. Use when the user…
paper-writer
Drafts publishable paper prose from the author's own materials, from a single paragraph to a full manuscript, across STEM and non-STEM fields. Every factual claim traces to user input, verified retrieval, or field common knowledge; citations pass an independent verification ladder; delivery is clean prose with zero…
pre-submission-reviewer
Runs a pre-submission review of a technical paper across five dimensions: macro logic, writing details, English grammar, LaTeX formatting, and figure quality. Uses a reviewer-style severity taxonomy (CRITICAL / MAJOR / MINOR) and flags banned AI-tone vocabulary and em-dash misuse. Use when the user asks to 'review…
benchmark-paper-template
Structures Benchmark and Evaluation papers using the five-pillar framework (Research Gap, Construction Pipeline, Evaluation Framework, Empirical Findings, optional Companion Method). Returns a completeness audit, a six-part Introduction logic chain, a Section 2-7 skeleton, and a pre-submission checklist. Use when…
deep-research
Runs a deep, survey-grade literature investigation on a research topic: freezes research questions, searches from multiple adversarial perspectives, verifies every citation, synthesizes evidence into a MECE taxonomy with in-sentence cross-comparison, and answers the research questions in an evidence-first report. Use…