content-extract

content-extract is a skill for Claude Code from aAAaqwq/AGI-Super-Team. It costs 92 tokens per session (1,041 once invoked), scanned A, original, MIT.

A URL-to-Markdown extraction workflow for turning web pages into readable Markdown. It first tries a lightweight fetch and uses the MinerU service when a page is blocked, incomplete, or difficult to parse.

In plain words
What is it for?
Use it to extract article content, especially from difficult sites such as WeChat pages, while retaining the source URL and paths to the extracted files.
Why use it?
It provides a fallback when normal web fetching returns an error, a login or anti-bot page, or only part of the article.

Skill for Claude Code

Written for Claude Code: shipped in a Claude Code plugin. Also seen: built for openclaw.

Needs its repository: it runs a file that does not travel with it, so clone the repository first. The line is python3 mineru-extract/scripts/mineru_parse_documents.py \.

Part of the agi-super-team plugin — 193 skills, 1 agent shipped together

Good fit Use it to extract article content, especially from difficult sites such as WeChat pages, while retaining the source URL and paths to the extracted files.

Compare 6 skills from other repositories ↓
Install

Getting it into your agent

It runs from inside its repository, so the clone comes first — what it calls does not travel with the file alone.

Clone the repo
git clone --depth 1 https://github.com/aAAaqwq/AGI-Super-Team
agentmods
npx agentmods add skills/aaaaqwq/agi-super-team/content-extract

Made for: Claude Code.

Or install agi-super-team, the plugin that ships this one along with the rest of its 193 skills, 1 agent.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for content-extract

README.md
[![agentmods](https://agentmods.dev/badge/skills/aaaaqwq/agi-super-team/content-extract/github.svg)](https://agentmods.dev/skills/aaaaqwq/agi-super-team/content-extract)
Your own site
<a href="https://agentmods.dev/skills/aaaaqwq/agi-super-team/content-extract"><img src="https://agentmods.dev/badge/skills/aaaaqwq/agi-super-team/content-extract/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for content-extract

Your own site · 80×15
<a href="https://agentmods.dev/skills/aaaaqwq/agi-super-team/content-extract"><img src="https://agentmods.dev/badge/skills/aaaaqwq/agi-super-team/content-extract.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 92 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,041 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • NVIDIA SkillSpector pass 7 Sept 2026
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00092 $0.01041
Opus 5 $0.00046 $0.00521
Sonnet 5 $0.00018 $0.00208
Haiku 4.5 $0.00009 $0.00104

Measured 4d ago against content hash 3b10f04812b4, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-08, from the pricing page.

Security

Grade A, and why

content-extract scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

The scan reads SKILL.md. This mod also ships 1 executable file (scripts/content_extract.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/content-extract/SKILL.md · 99 lines

How it starts

The opening of the file, as written. The whole thing — 99 lines — stays where its author put it; the contents beside it link to each section on GitHub.

content-extract — 上层内容解析入口(MCP 语义对齐,但不跑 MCP Server)

  • Author: Daniel Li
  • Copyright © Daniel Li. All rights reserved.

目标:把“给我一个 URL → 产出可读 Markdown + 可追溯入口”变成一个统一入口,供后续所有业务 skill(github-explorer、写作类 skills、日报等)复用。

核心原则(来自你发的 Excel Skill 拆解文章的启发):

  • 行为规约层:永远给出可追溯入口(原文 URL + 解析产物路径/链接),绝不编造来源。
  • Token 探针:先用低成本 probe 判断可不可以直接抓;不行再走重解析(MinerU)。
  • 反弹机制:失败时返回“下一步动作建议”,而不是一堆异常栈。

工作流(Decision Tree)

输入:url

  1. Domain Whitelist(跳过 probe):若 URL 属于高概率反爬/动态站点(微信/知乎等),直接走 MinerU
  • 白名单文件:references/domain-whitelist.md
  • 对命中白名单的 URL:强制 model_version=MinerU-HTML
  1. Probe(低成本):优先用 web_fetch(url)
  • 目标:拿到正文 markdown(便宜、快)
  • 判断“失败/不合格”条件(见 references/heuristics.md)包括:
    • 403/401/反爬
    • 只有“环境异常/验证码/请在微信打开”等提示
    • 内容极短/明显导航页/丢正文
  1. Fallback(高保真):走 MinerU 官方 API
  • 调用下游 driver:skills/mineru-extract/scripts/mineru_parse_documents.py
  • 对 HTML 页面(微信等):强制 model_version=MinerU-HTML
  1. 输出统一结果合同(Result Contract)

无论用 probe 还是 MinerU,都返回同一套结构:

{
  "ok": true,
  "source_url": "...",
  "engine": "web_fetch" ,
  "markdown": "...",
  "artifacts": {
    "out_dir": "...",
    "markdown_path": "...",
    "zip_path": "..."
  },
  "sources": [
    "原文URL",
    "(如使用MinerU)MinerU full_zip_url",
    "(如使用MinerU)本地markdown_path"
  ],
  "notes": ["任何重要限制/失败原因/下一步建议"]
}

注意:engine 可能是 web_fetchmineru

MinerU 调用(给 agent 的确定性脚本)

当需要 MinerU 时,用这个命令(返回 JSON,且可把 markdown 内联进 JSON,便于下游总结):

python3 mineru-extract/scripts/mineru_parse_documents.py \
  --file-sources "<URL>" \
  --model-version MinerU-HTML \
  --emit-markdown --max-chars 20000

路径说明: 上述命令假设你在 skills 安装根目录下执行。如果 mineru-extract 安装在其他位置,请替换为实际路径。

交付规范(强制)

  • 输出必须包含 sources(原文入口 + 解析产物入口)。
  • 如果 MinerU 成功:必须把 markdown_path(本地路径)写进 sources,方便复查。
  • 如果两条链路都失败:必须明确失败原因,并给出下一步(例如:让 Boss 提供可访问镜像链接 / 允许我用浏览器 relay 导出 HTML / 走上传 HTML 文件解析的兜底方案)。

本 skill 自身不做什么

  • 不跑 MCP Server(避免常驻服务与运维负担)
  • 不试图绕过登录/验证码(这属于访问层问题;我们只做解析层和工作流路由)

References

Read the full file on GitHub · 99 lines

Files

What ships with it

3 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 99 lines · 92 tokens per session scan A 3b10f04812b4

Subscribe to this mod's changes

content-extract is a skill published in the GitHub repository aAAaqwq/AGI-Super-Team (91 stars, last pushed yesterday), licensed MIT. It adds 92 tokens to every session and 1,041 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-05.

Related

Other skills, from other repositories

deslop

The optimization pass, defined - delete before you add, one smell class per pass, behaviour pinned by a test that ran BEFORE the edit. Lints a SKILL.md and prose by the same instinct. Use for the per-story optimization pass or when code has grown noisy without growing capable.

jjanczur/tyran · 58 tokens

root-cause

Find the mechanism behind a failure instead of patching its symptom - reproduce first, one variable per experiment with the prediction written before the run, exit by naming the mechanism and pinning it with a failing test. Use for a bug, an unexplained red test, or a failure that will not reproduce.

jjanczur/tyran · 61 tokens

guidance

Add, edit, or audit guidance docs. Default writes guidance for Claude (.claude/guidance/, Markdown, moflo universal rules). -h writes for human readers (docs/, lighter ruleset). --html emits HTML with a minimal default stylesheet instead of Markdown. -a audits the .claude/guidance/ directory.

eric-cielo/moflo · 70 tokens

eldar

Consult the Eldar — audit a project's moflo + Claude Code setup for portable, high-leverage gaps and guide remediation. Default mode is read-only audit with severity-ranked findings; --fix presents an interactive triage menu and walks the user through each chosen fix (healer, missing CLAUDE.md, sparse guidance…

eric-cielo/moflo · 115 tokens

aigon-next

Suggest the most likely next workflow action based on current context.

jayvee/aigon · 15 tokens

review-deep

Drive the deep-review phase of an automated PR review. Consumes the walkthrough, runs the deterministic deep-review workflow (parallel lenses → adversarial validation → code-enforced threshold/caps), drafts the surviving findings, and completes the review run.

ShreyPaharia/octomux · 52 tokens