Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add AkaliKong/PaperClaw --skill paper-deep-parsergit clone --depth 1 https://github.com/AkaliKong/PaperClawWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/akalikong/paperclaw/paper-deep-parser)<a href="https://agentmods.dev/skills/akalikong/paperclaw/paper-deep-parser"><img src="https://agentmods.dev/badge/skills/akalikong/paperclaw/paper-deep-parser.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00060 | $0.01001 |
| Opus 5 | $0.00030 | $0.00500 |
| Sonnet 5 | $0.00012 | $0.00200 |
| Haiku 4.5 | $0.00006 | $0.00100 |
Grade A, and why
paper-deep-parser scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 104 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Paper Deep Parser — Knowledge Card Structured Extraction
你负责对最终选中的论文进行深度精读,并从知识卡片中提取结构化信息。
⚠️ 关键规则
- 每篇论文使用独立 Session:通过
sessions_spawn创建,避免上下文污染。 - 先精读后解析:先触发
read-arxiv-paper生成 card.md,再调用card_parser.py提取结构化字段。 - N/A 降级:无法提取的字段填入
"N/A",不要猜测或编造。
执行流程
Step 1: 获取待精读论文列表
读取 pipeline_data/{run_id}/skill3_final_selection.json,获取所有需要精读的论文。
Step 2: 逐篇触发精读
对每篇论文:
-
检查是否已有 card.md:
- 搜索
research/papers/下是否有对应的 card.md - 如果已有,跳过精读,直接进入 Step 3
- 搜索
-
创建独立 Session 精读:
sessions_spawn: 创建新 Session sessions_send: 触发 read-arxiv-paper Skill,输入论文 arXiv URL session_status: 轮询直到完成 -
并发控制:同时运行的精读 Session 不超过 3 个,避免资源竞争。
Step 3: 提取结构化字段
对每篇已有 card.md 的论文,调用解析脚本:
python $PAPER_AGENT_ROOT/scripts/card_parser.py \
--card-path /path/to/card.md \
--arxiv-id {arxiv_id}
或使用批量模式(自动查找所有 card.md):
python $PAPER_AGENT_ROOT/scripts/card_parser.py --run-id {run_id}
Step 4: 检查结果
解析脚本输出 JSON 到 stdout,同时保存到 pipeline_data/{run_id}/skill4_parsed/{arxiv_id}.json。
检查 parse_success 和 fields_extracted 字段:
parse_success: true+fields_extracted >= 3= 解析良好parse_success: true+fields_extracted < 3= 部分解析,可能需要 card.md 格式优化parse_success: false= 解析失败,检查parse_errorneeds_reading: true= 缺少 card.md,需要先触发 read-arxiv-paper
提取字段说明
| 字段 | 含义 | 示例 |
|---|---|---|
sub_field |
研究子领域 | generative_rec, sequential_rec |
ID_paradigm |
ID/表示范式 | Semantic ID, RQ-VAE, Collaborative ID |
item_tokenizer |
物品标记化方法 | RQ-VAE, BPE, SentencePiece |
baselines_compared |
对比的基线方法 | ["SASRec", "BPR", "BERT4Rec"] |
transferable_techniques |
可迁移的技术 | ["Semantic ID generation", "Multi-task loss"] |
inspiration_ideas |
启发的研究想法 | ["Combine X with Y for Z"] |
输出文件
每篇论文输出到 pipeline_data/{run_id}/skill4_parsed/{arxiv_id}.json:
{
"arxiv_id": "2305.XXXXX",
"title": "Example Paper Title...",
"sub_field": "your_research_field",
"ID_paradigm": "Semantic ID",
"item_tokenizer": "Tokenizer Method",
"baselines_compared": ["SASRec", "BPR", "BERT4Rec"],
"transferable_techniques": ["Technique A from the paper"],
"inspiration_ideas": ["Combine X with Y for Z"],
"card_path": "$PAPER_AGENT_ROOT/research/papers/2305.XXXXX_example/card.md",
"parse_success": true,
"fields_extracted": 6,
"fields_total": 6
}
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 104 lines · 60 tokens per session scan A 0c5a8e98e003
paper-deep-parser is a skill published in the GitHub repository AkaliKong/PaperClaw (22 stars, last pushed 6mo ago), licensed MIT. It adds 60 tokens to every session and 1,001 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
local-ai-agents
Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
next-cache-components-optimizer
Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…
next-partial-prefetching-adoption
Turn on Partial Prefetching in a Next.js app and work through the insights it surfaces. Use when the user wants to enable or adopt Partial Prefetching, flip the partialPrefetching flag, opt routes in with export const prefetch = 'partial', audit Link prefetch={true} behavior, preserve existing prefetched UI with…
chronicle
Analyze Copilot session history for standup reports, usage tips, session search, and session reindexing. Use when the user asks for a standup, daily summary, usage tips, workflow recommendations, wants to search or find past sessions by keyword/file/PR, wants to reindex their session store, or asks about deleting…