scrape-strategy

scrape-strategy is a skill for Claude Code, Codex from blisspixel/primr. It costs 39 tokens per session (362 once invoked), scanned A, original, Apache-2.0.

A guide for choosing how to proceed when a website is difficult to scrape, meaning software cannot reliably retrieve its pages. It compares research modes and recommends retrying, changing mode, or accepting partial coverage.

In plain words
What is it for?
Use it when a protected or poorly accessible website produces weak results, to review the current run, choose a suitable research mode, and set realistic coverage expectations.
Why use it?
It keeps site-access problems from turning into unnecessary troubleshooting. The guide helps decide whether to rely on the site or switch to broader external research.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/blisspixel/primr/scrape-strategy
Any agent
npx skills add blisspixel/primr --skill scrape-strategy
Clone the repo
git clone --depth 1 https://github.com/blisspixel/primr

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for scrape-strategy

README.md
[![agentmods](https://agentmods.dev/badge/skills/blisspixel/primr/scrape-strategy.svg)](https://agentmods.dev/skills/blisspixel/primr/scrape-strategy)
Your own site
<a href="https://agentmods.dev/skills/blisspixel/primr/scrape-strategy"><img src="https://agentmods.dev/badge/skills/blisspixel/primr/scrape-strategy.svg" alt="Measured on agentmods" height="20"></a>
Per session 39 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 362 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00039 $0.00362
Opus 5 $0.00019 $0.00181
Sonnet 5 $0.00008 $0.00072
Haiku 4.5 $0.00004 $0.00036

Measured 4d ago against content hash 266a9dcd0613, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

scrape-strategy scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/scrape-strategy/SKILL.md · 44 lines

What it actually says

Scrape Strategy

Purpose

Use this skill when the bottleneck is site access rather than report writing. Focus on deciding whether to stay in scrape-capable modes or pivot to external research.

Workflow

  1. Read primr://research/modes to compare scrape, deep, full, and premium behavior.
  2. Review current run state through primr://research/status or check_jobs.
  3. Use references/tiers.md only when you need tier-level scraping context.
  4. Recommend the smallest viable change: retry, change mode, or accept partial coverage.

Decision Rules

  • Prefer deep when the target site is heavily protected or first-party signal is sparse.
  • Stay with scrape or full when first-party pages are the core evidence source.
  • Do not promise exact success percentages unless the run data already shows them.
  • Escalate from site troubleshooting to mode selection quickly; avoid over-explaining scrape internals unless the user asks.

Example

User: The site seems blocked

1. Read primr://research/status
2. Read primr://research/modes
3. Explain whether deep mode is the better fit
4. If needed, estimate_run(company_url="https://example.com", mode="deep")
Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 44 lines · 39 tokens per session scan A 266a9dcd0613

Subscribe to this mod's changes

scrape-strategy is a skill published in the GitHub repository blisspixel/primr (3 stars, last pushed yesterday), licensed Apache-2.0. It adds 39 tokens to every session and 362 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

vaaya

Vaaya is the payment system for agents: one MCP server that lets your agent call paid APIs pay-per-call with no API keys. Web search, scraping, image and video generation, code sandboxes, browser automation, email, phone calls, lead enrichment, live data. Priced in cents per call, billed only on success, and every…

vaaya-ai/vaaya-mcp · 0 tokens

fetch-crawl4ai

Deep URL fetch using crawl4ai (Playwright-powered) for JS-rendered pages, anti-bot sites, and dynamic content. Slower than fetch-jina but handles sites that block simple fetchers. Requires user-installed crawl4ai package.

0xmariowu/Autosearch · 56 tokens

chrome-devtools-cli

Use this skill to write shell scripts or run shell commands to automate tasks in the browser or otherwise use Chrome DevTools via CLI.

ChromeDevTools/chrome-devtools-mcp · 31 tokens

chrome-devtools

Uses Chrome DevTools via MCP for efficient debugging, troubleshooting and browser automation. Use when debugging web pages, automating browser interactions, analyzing performance, or inspecting network requests. This skill does not apply to --slim mode (MCP configuration).

ChromeDevTools/chrome-devtools-mcp · 55 tokens

2-repro-issue

Reproduce a single LinkedIn-MCP issue locally on the current branch against the real authenticated LinkedIn session at /.linkedin-mcp/profile/, using the MCP streamable-http server. Captures the exact failure mode (tool output, error, missing data) and maps it back to the scraper code path. Use when the user says…

stickerdaniel/linkedin-mcp-server · 75 tokens

ddgs

Use when an agent needs to search the web, find images/news/videos/books, or extract content from a URL. Covers the ddgs Python library, CLI, and MCP server integration.

deedy5/ddgs · 40 tokens