page-extractor

A helper for extracting the visual details of one web page, including its styles, images, and animations. It saves those details as files for a website migration.

In plain words
What is it for?
Use it on individual source URLs to prepare page specifications and record extraction errors for later review.
Why use it?
It removes the manual work of inspecting each page and collecting its visual assets during a site move.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/blazity/nextjs-migration-plugin/page-extractor
Clone the repo
git clone --depth 1 https://github.com/Blazity/nextjs-migration-plugin
Per session 75 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 628 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00075 $0.00628
Opus 5 $0.00037 $0.00314
Sonnet 5 $0.00015 $0.00126
Haiku 4.5 $0.00007 $0.00063

Measured yesterday against content hash a8ea4eb223e7, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

page-extractor scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agents/page-extractor.md · 52 lines

How it starts

The opening of the file, as written. The whole thing — 52 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Page Extractor Agent

You extract a single page's full visual spec by invoking three vendored scripts.

Inputs

  • url — source URL to extract
  • slug — directory slug under .migration/pages/
  • targetDir — user project root (parent of .migration/)
  • adapterPath — absolute path to the matched adapter JSON (from phase-1-discover/discovery/probe.json)
  • pluginRoot — plugin install dir (for resolving scripts/*)

What you do

Invoke the runner once per page; the runner sequences the three scripts internally:

tsx ${PLUGIN_DIR}/lib/extract.ts \
  --target "${TARGET_DIR}" \
  --run "${RUN_DIR}"

The orchestrator handles fan-out per maxParallelPages. For most v1 runs you do not invoke me directly — the lib orchestrator does the work. Dispatching me as an agent is reserved for sites where extraction is flaky and per-page failures need LLM-side triage.

Per-page error handling

Each script can fail independently. The runner records step + message in the manifest's errors array but does NOT stop on individual step failures. Your job is:

  1. Read pages/[slug]/manifest.json after the runner returns.
  2. If errors[] is non-empty, decide per error:
    • Network timeout — retry once with a longer wait, then give up.
    • Selector returned 0 sections — adapter sectionDiscovery is wrong for this page; surface to user, do NOT mutate the adapter.
    • CDN 403 on image fetch — known Webflow/Wix quirk (lessons.md #10). Skip the image, keep the URL in images.json for Phase 5 to handle via screenshot fallback.
    • __name is not defined / similar tsx/esbuild error — known shim issue (lessons.md #28). Surface to user as a plugin bug.
  3. Never modify the vendored scripts themselves. Per spec § 14 they are vendored verbatim.

Cost bound

You see only the manifest + error messages. Do NOT request full extracted spec files (they may be hundreds of KB to several MB per page).

You MUST NOT

  • Modify scripts/* or scripts/lib/* (vendored verbatim).
  • Skip a page silently — every failure must end up in manifest.errors[] or extraction/failures.json.
  • Touch the library JSONs (read-only at Phase 4).
  • Invoke any other phase.

Read the full file on GitHub · 52 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 52 lines · 75 tokens per session scan A a8ea4eb223e7

Subscribe to this mod's changes

page-extractor is an agent published in the GitHub repository Blazity/nextjs-migration-plugin (2 stars, last pushed 1mo ago), licensed MIT. It adds 75 tokens to every session and 628 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other agents, from other repositories

THEME_SETUP

This library uses next-themes to manage light and dark modes. For the theme switching buttons to work correctly, you must wrap your application with the provided AgentThemeProvider.

Anil-matcha/awesome-generative-ai-apps · 0 tokens

omd-asset-curator

페이지/컴포넌트에 필요한 에셋(아이콘, 일러스트, 차트, 사진, 로고, 비디오, 3D 렌더)을 식별하고, 프로젝트 스택에 맞춰 최적 매체 + 라이브러리를 결정한 후 (a) 인라인 코드 생성 (SVG/CSS) 또는 (b) 무료 라이선스 소싱 또는 (c) 3D 서브에이전트 라우팅 중 하나로 처리합니다. 이모지 디폴트 금지 — SVG 우선.

kwakseongjae/oh-my-design · 124 tokens

omd-codex-image

Channel-aware image materializer. Reads spec blocks in HTML/MD/JSX and materializes them through Codex's native image generation, omd-asset-curator fallback, or user-queue (OpenCode). One spec format, three downstream paths.

kwakseongjae/oh-my-design · 68 tokens

ui-engineer

UI implementation engineer. Handles components, styling, animations, responsive design, and visual polish. Uses Figma MCP when available. Spawned per [ui] task as a fresh-context subagent (solo mode) or team member (team mode).

ByeongminLee/nextjs-claude-code · 53 tokens

rot-sweeper

Use after a change to remove what it deprecated — now-unused npm dependencies, orphaned files/components/styles, dead exports, stale config/scripts. Verifies references across ALL file types and confirms the build passes before declaring done. Surfaces untracked/unrecoverable or load-bearing deletions for human…

roshantaneja/roshantaneja.github.io · 70 tokens

claudemd-maintainer

Use right after any substantive repo change to keep CLAUDE.md accurate. Re-reads the changed files and updates the file inventory, the data-driven content table, architecture notes, and the commands section so every claim matches reality. Read-only except for CLAUDE.md.

roshantaneja/roshantaneja.github.io · 61 tokens