Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add AkaliKong/PaperClaw --skill paper-repo-evaluatorgit clone --depth 1 https://github.com/AkaliKong/PaperClawWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/akalikong/paperclaw/paper-repo-evaluator)<a href="https://agentmods.dev/skills/akalikong/paperclaw/paper-repo-evaluator"><img src="https://agentmods.dev/badge/skills/akalikong/paperclaw/paper-repo-evaluator/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/akalikong/paperclaw/paper-repo-evaluator"><img src="https://agentmods.dev/badge/skills/akalikong/paperclaw/paper-repo-evaluator.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00051 | $0.00726 |
| Opus 5 | $0.00026 | $0.00363 |
| Sonnet 5 | $0.00010 | $0.00145 |
| Haiku 4.5 | $0.00005 | $0.00073 |
Grade A, and why
paper-repo-evaluator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Paper Repo Evaluator — Code Repository Assessment
你负责评估最终选中论文的代码仓库质量和可集成性。
⚠️ 关键规则
- 三阶段搜索:先从论文文本提取链接 → 找不到则搜索 GitHub → 完全未找到标记
has_code: false。 - API 失败降级:GitHub API 不可用时标记
github_api_failed: true,不要阻塞后续流程。 - 尊重 Rate Limit:GitHub API 调用间隔至少 2 秒,遇到 403 时指数退避重试。
执行流程
Step 1: 调用评估脚本
使用批量模式评估当前 run 的所有选中论文:
python $PAPER_AGENT_ROOT/scripts/repo_evaluator.py --run-id {run_id}
或评估单篇论文:
python $PAPER_AGENT_ROOT/scripts/repo_evaluator.py \
--arxiv-id {arxiv_id} \
--title "Paper Title"
Step 2: 检查结果
脚本返回 JSON 统计:
{
"step": "repo-eval",
"status": "success",
"total": 5,
"has_code_count": 3,
"no_code_count": 2,
"api_failed_count": 0
}
Step 3: 向用户报告
向用户报告代码评估结果摘要:
- X/Y 篇论文有关联代码
- 主要编程语言分布
- 集成成本评估
评估维度
| 维度 | 说明 |
|---|---|
has_code |
是否找到代码仓库 |
github_url |
仓库 URL |
stars |
GitHub Stars 数 |
language |
主要编程语言 |
integration_cost |
集成成本(Low/Medium/High) |
github_api_failed |
API 是否失败 |
代码链接搜索策略
- 文本提取:从论文摘要、card.md 中用正则提取 GitHub/GitLab 链接
- GitHub 搜索:如果文本中无链接,用论文标题搜索 GitHub
- 标记未找到:两种方式都失败时,标记
has_code: false
输出文件
每篇论文输出到 pipeline_data/{run_id}/skill5_repo_eval/{arxiv_id}.json:
{
"arxiv_id": "2305.05065",
"has_code": true,
"github_url": "https://github.com/owner/repo",
"platform": "github",
"stars": 256,
"forks": 42,
"language": "Python",
"license": "MIT",
"integration_cost": "Low",
"github_api_failed": false,
"search_method": "extracted_from_text"
}
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 102 lines · 51 tokens per session scan A ced96d185d48
paper-repo-evaluator is a skill published in the GitHub repository AkaliKong/PaperClaw (22 stars, last pushed 6mo ago), licensed MIT. It adds 51 tokens to every session and 726 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
local-ai-agents
Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…
insight-error-page
Write or audit an insight-kind error page for the Next.js dev overlay. Use when creating a new errors/ .mdx page, auditing an existing one, or checking that a page matches the framework fix cards. Covers page structure, title alignment, FixCard cards with Copy prompt button, code snippets, terminology verification…