Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/scraperapi/scraperapi-skills/scraperapi-clinpx skills add scraperapi/scraperapi-skills --skill scraperapi-cligit clone --depth 1 https://github.com/scraperapi/scraperapi-skillsWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00262 | $0.03549 |
| Opus 5 | $0.00131 | $0.01775 |
| Sonnet 5 | $0.00052 | $0.00710 |
| Haiku 4.5 | $0.00026 | $0.00355 |
Grade A, and why
scraperapi-cli scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 316 lines — stays where its author put it; the contents beside it link to each section on GitHub.
ScraperAPI CLI (sapi)
sapi is the official ScraperAPI command-line tool. It is the right choice when:
- The user is already in a terminal and wants a result now without writing a script.
- A scrape is part of a shell pipeline (
sapi … | jq …,xargs,make, GitHub Actions). - A one-off scheduled task (cron, launchd, systemd timer) needs to hit ScraperAPI without a project setup.
- The user is exploring — testing whether
--renderor--premiumunblocks a target before committing the choice to code.
If the user is writing application code in Python, Node, PHP, Ruby, or Java, point them at the matching SDK skill instead — the CLI is for shells, not application logic.
Install and authenticate
npm install -g scraperapi-cli # requires Node.js 18+
sapi init # interactive: prompts for the key and validates it
Non-interactive setup (for CI / Dockerfiles):
sapi init --api-key "$SCRAPERAPI_API_KEY"
Key resolution order
sapi looks for the API key in this order, stopping at the first hit:
--api-key <key>flag on the commandSCRAPERAPI_API_KEYenvironment variable~/.config/scraperapi/config.json(written bysapi init)
In CI, prefer the env var — it keeps the key out of shell history and config files.
Output contract — important for piping
| Stream | What goes there |
|---|---|
| stdout | The data (page body, JSON, table rows) |
| stderr | Spinners, warnings, errors |
When stdout is not a TTY (a pipe or redirect), sapi automatically switches to JSON mode:
sapi scrape https://example.com | jq .body # auto-JSON
sapi scrape https://example.com > out.json # auto-JSON
sapi scrape https://example.com # human output (page body to stdout)
Force JSON mode in a TTY with --json. Force a specific body format with --output html|markdown|text|json|csv — --output overrides the non-TTY auto-JSON rule, so you can pipe raw HTML or markdown into another tool without it being wrapped in JSON.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 316 lines · 262 tokens per session scan A 1b404b9771cc
scraperapi-cli is a skill published in the GitHub repository scraperapi/scraperapi-skills (10 stars, last pushed 25d ago), licensed MIT. It adds 262 tokens to every session and 3,549 once invoked, about $0.0013 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
scrapling-official
Scrape web pages using Scrapling with anti-bot bypass (like Cloudflare Turnstile), stealth headless browsing, spiders framework, adaptive scraping, and JavaScript rendering. Use when asked to scrape, crawl, or extract data from websites; webfetch fails; the site has anti-bot protections; write Python code to…
ketch
Research skill for ketch — a fast stateless CLI for web search, OSS code search, curated library docs, page scraping, and site crawling; an optional MCP server exists for operators who want it, but the CLI is the primary interface. Use when a question needs live sources: 'research X', 'what are people saying about Y'…
browser-automation
Playwright-based browser automation patterns for autonomous web interaction.
web-browse
Drive a headless browser to navigate pages, read content, click, and fill forms — for sites that need JavaScript rendering or interaction beyond a plain HTTP fetch.
ecommerce-full-pipeline
电商运营在开展跨境电商或闲鱼捡漏业务时,若需解决选品难、上架繁琐等痛点,必用此技能!一键打通“爆品挖掘→1688采集→多平台上架→推广文案→短视频生成”全自动流水线,轻松实现端到端自动化,让开店运营效率翻倍。.
browser-scrape
用 AutoCLI 二进制驱动用户已登录的 Chrome 抓取 Twitter/X、知乎、Bilibili、Reddit 等 55+ 站点.