Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/secondlifes/code-intel/rad-web-scrapingnpx skills add SecondLifes/code-intel --skill rad-web-scrapinggit clone --depth 1 https://github.com/SecondLifes/code-intelWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00095 | $0.00977 |
| Opus 5 | $0.00048 | $0.00489 |
| Sonnet 5 | $0.00019 | $0.00195 |
| Haiku 4.5 | $0.00010 | $0.00098 |
Grade A, and why
rad-web-scraping scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
This is a copy
100% identical to rad-web-scraping — 0 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
How it starts
The opening of the file, as written. The whole thing — 39 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Web Scraping / Structured Data Extraction
When to Use
- Pulling structured data (prices, listings, articles, tables) from a website for a script, pipeline, or dataset.
- Turning arbitrary web pages into clean Markdown/text for an LLM, RAG index, or agent to consume.
- Deciding between a plain HTTP request, a scraping library, and full browser automation for a given target.
- Reviewing or debugging an existing scraper that's fragile or getting blocked.
Usage
| You say | What happens |
|---|---|
"Pull the prices from this site" / "Bu siteden ilan verilerini çek" |
Discovery-priority check first (API → embedded JSON → sitemap/feed → DOM as last resort, full order in references/discovery-and-extraction-patterns.md), then references/tool-selection.md's decision tree picks the library. |
"Turn these pages into clean Markdown for my RAG index" |
Recommends crawl4AI (LLM/RAG-ready output) per references/tool-selection.md. |
"My scraper keeps getting blocked / breaks randomly" |
Reviews against the retry/backoff/rate-limit baseline in references/discovery-and-extraction-patterns.md — a scraper missing these isn't hardened, it's incomplete. |
"Bypass this CAPTCHA / get around their bot detection" |
Refused — out of scope regardless of phrasing; reasonable retries/rate-limiting are fine, detection evasion is not. |
A scraping request with no target site/URL named ("Scrape it", "veriyi çek") |
Asks what site/data is meant before picking any tool — never guesses a target, and an empty request is never treated as "nothing to do." |
Golden Rules
- Discover before you parse. A REST/GraphQL API, a JSON blob embedded in the page's HTML, or a sitemap/feed is more stable than any CSS selector — always look for one before writing DOM-parsing code. Full priority order:
references/discovery-and-extraction-patterns.md. - Plain HTTP before a browser. Reach for
httpx/requestsfirst; only escalate to Playwright/browser automation when the target genuinely requires JS execution or session state a plain client can't replicate. A headless browser is slower, heavier, and more fragile than it looks. - Pick the tool for the job, not the first one you remember —
references/tool-selection.mdhas a verified decision tree (crawl4AI vs ScrapeGraphAI vs raw HTTP vs Playwright) with install commands, license, and when each wins. - Retry, back off, and rate-limit by default. A scraper with no retry/backoff either gets blocked or silently drops data the first time the network hiccups — this isn't optional hardening, it's baseline correctness for anything hitting a live site repeatedly.
- Scope and ethics — non-negotiable. Only scrape data you're legitimately allowed to access; respect
robots.txtand a site's terms of service. Never write code whose purpose is bypassing CAPTCHAs, defeating bot-detection, or evading paywalls/access controls — that's out of scope for this skill regardless of what's requested, matching this environment's own prohibition on bot-detection bypass. Reasonable technical robustness (retries, realistic headers, rate limiting) is fine; detection evasion is not.
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 39 lines · 95 tokens per session scan A f0911fdc353e
rad-web-scraping is a skill published in the GitHub repository SecondLifes/code-intel (2 stars, last pushed 21d ago), licensed Apache-2.0. It adds 95 tokens to every session and 977 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. It is 100% identical to rad-web-scraping, differing in 0 lines, and is treated as a copy.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
babysit-pr
Babysit a GitHub pull request after creation by continuously polling review comments, CI checks/workflow runs, and mergeability state until the PR is merged/closed or user help is required. Diagnose failures, retry likely flaky failures up to 3 times, auto-fix/push branch-related issues when appropriate, and keep…
imagegen
Generate or edit raster images when the task benefits from AI-created bitmap visuals such as photos, illustrations, textures, sprites, mockups, or transparent-background cutouts. Use when Codex should create a brand-new image, transform an existing image, or derive visual variants from references, and the output…
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…
next-cache-components-optimizer
Drive a Next.js route to instant navigation by setting up an agentic loop, under Cache Components / PPR, on initial load (hard navigation) and client-side navigation (soft navigation). Encode the goal as a failing @next/playwright instant() e2e and work it to green, one verified route at a time; the shipped test then…