Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add instructions/collaborative-deep-research/agent-papers-cli/claude-mdgit clone --depth 1 https://github.com/collaborative-deep-research/agent-papers-cliWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00909 | $0.00909 |
| Opus 5 | $0.00454 | $0.00454 |
| Sonnet 5 | $0.00182 | $0.00182 |
| Haiku 4.5 | $0.00091 | $0.00091 |
Grade A, and why
agent-papers-cli CLAUDE.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 39 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Project: paper & paper-search CLI
Two CLI tools in one repo: paper (read academic PDFs) and paper-search (web + academic search). See README.md for full docs and SKILLS.md for agent workflows.
Quick reference
- Entry points:
paper = paper.cli:cli,paper-search = search.cli:cli(Click) - paper modules:
cli.py,parser.py,fetcher.py,storage.py,renderer.py,models.py,highlighter.py,layout.py,bibtex.py - search modules:
cli.py,config.py,models.py,renderer.py,backends/{google,semanticscholar,pubmed,browse}.py - Cache:
~/.papers/<paper_id>/(papers:paper.pdf,parsed.json,metadata.json,highlights.json,layout.json,layout/*.png,paper_annotated.pdf,bibtex.bib),~/.papers/.models/(YOLO weights),~/.papers/.env(persistent API keys),~/.papers/.last_header(header auto-suppression state) - Local PDFs: Pass a file path (e.g.,
./paper.pdf) instead of an arxiv ID — reads directly, no download. Cache uses{stem}-{hash8}IDs (SHA-256 of absolute path) to avoid collisions. Stale caches are detected via mtime comparison. - Tests:
pytest— paper tests intests/(124 tests), search tests intests/search/(69 tests) - Agent skills:
.claude/skills/— research-coordinator, deep-research, literature-review, fact-check
Architecture notes
paper
- Parser tries PDF built-in outline first, falls back to font-size heuristics
- GROBID backend planned for future (noted in
parser.py) - Data model inspired by papermage: flat
raw_text+Sectionlist with character-offsetSpans - Downloads use atomic temp-file-then-rename pattern
- Storage sanitizes paper IDs to prevent path traversal
- Layout detection (optional
[layout]extra) uses DocLayout-YOLO via doclayout_yolo - Layout detection is lazy: runs on first
paper figures/tables/equationscall, cached inlayout.json - Supports MPS (Apple Metal), CUDA, and CPU backends for inference
- Model weights auto-downloaded from collab-dr/DocLayout-YOLO-DocStructBench (pinned fork) to
~/.papers/.models/on first use - Detected elements are cropped as PNG screenshots to
~/.papers/<id>/layout/
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 39 lines · 909 tokens per session scan A 240aa74c5dfc
agent-papers-cli CLAUDE.md is an instructions file published in the GitHub repository collaborative-deep-research/agent-papers-cli (50 stars, last pushed 5mo ago), licensed Apache-2.0. It adds 909 tokens to every session, about $0.0045 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other instructions, from other repositories
vscode buildNext.instructions.md
Working notes and architecture documentation for the new esbuild-based build system in build/next. Use when making changes to the new build pipeline (transpile/bundle commands, NLS plugin, source-map handling, resource copying, or self-hosting watch tasks).
spec-kit AGENTS.md
AGENTS.md instructions for github/spec-kit, covering agents.md, about spec kit and specify, quickstart — add a new integration in 5 steps, integration architecture and integrationmanifest — file tracking.
codex AGENTS.md
AGENTS.md instructions for openai/codex, covering rust/codex-rs, the codex-core crate, code review rules, crate api surface and model visible context.
langchain AGENTS.md
AGENTS.md instructions for langchain-ai/langchain, covering global development guidelines for the langchain monorepo, corridor security analysis, project architecture and context, monorepo structure and development tools & commands.
vscode oss-third-party-notices.instructions.md
Instructions for microsoft/vscode, covering vs code oss third-party-notices pipeline, architecture, pipeline flow in ci, applying the notice (cutover) and fallback chain (never fail the build).
next.js AGENTS.md
Instructions for vercel/next.js, covering next.js development guide, codebase structure, monorepo overview, core package: packages/next and other important packages.