Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add mysleekdesigns/crawlforge-mcp --skill crawlforge-batch-automationgit clone --depth 1 https://github.com/mysleekdesigns/crawlforge-mcpWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/mysleekdesigns/crawlforge-mcp/crawlforge-batch-automation)<a href="https://agentmods.dev/skills/mysleekdesigns/crawlforge-mcp/crawlforge-batch-automation"><img src="https://agentmods.dev/badge/skills/mysleekdesigns/crawlforge-mcp/crawlforge-batch-automation/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/mysleekdesigns/crawlforge-mcp/crawlforge-batch-automation"><img src="https://agentmods.dev/badge/skills/mysleekdesigns/crawlforge-mcp/crawlforge-batch-automation.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00132 | $0.01406 |
| Opus 5 | $0.00066 | $0.00703 |
| Sonnet 5 | $0.00026 | $0.00281 |
| Haiku 4.5 | $0.00013 | $0.00141 |
Grade A, and why
crawlforge-batch-automation scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 130 lines — stays where its author put it; the contents beside it link to each section on GitHub.
CrawlForge Batch & Automation
Scale up scraping and drive interactive pages. Use batch_scrape for many URLs,
get_batch_results to page through async output, scrape_with_actions to
interact before scraping, and generate_llms_txt to produce a site's AI policy
file.
When to use
- "Scrape these 30 URLs" / "batch-scrape this list" →
batch_scrape - "Collect dozens of product / news / competitor pages" →
batch_scrape(async) - "Get the rest of the results from that batch" →
get_batch_results - "Click / type / scroll / wait before scraping" / "log in then extract" →
scrape_with_actions - "Generate an llms.txt for this site" →
generate_llms_txt
batch_scrape — many URLs in parallel (cost: 5)
Sync mode (results returned immediately), good for up to ~25 URLs:
{
"tool": "batch_scrape",
"params": {
"urls": ["https://a.com", "https://b.com", "https://c.com"],
"formats": ["markdown"],
"mode": "sync",
"maxConcurrency": 5
}
}
Async mode with a webhook for large batches (returns a batchId immediately):
{
"tool": "batch_scrape",
"params": {
"urls": ["https://a.com", "https://b.com"],
"formats": ["json"],
"mode": "async",
"webhook": { "url": "https://my-site.com/hook", "events": ["batch_completed", "batch_failed"] }
}
}
urlsaccepts plain strings OR objects{url, selectors, headers, timeout, metadata}for per-URL config. 1–50 URLs.formats:markdown,html,json,text.extractionSchemaapplies structured extraction to every URL.maxConcurrency1–20 (default 10);delayBetweenRequeststhrottles.- Sync batches over ~25 URLs trigger a confirmation prompt (elicitation) — use async for large jobs.
CLI: crawlforge batch urls.txt --format markdown --concurrency 10.
get_batch_results — page through output (cost: 1)
{ "tool": "get_batch_results", "params": { "batchId": "batch_1234567890_abc", "page": 2, "pageSize": 25 } }
Use the batchId from batch_scrape to retrieve paginated results for a
completed or in-progress job. Cheap (1 credit) because the batch was already
paid for. Completed jobs are also exposed as crawlforge://job/{jobId}
resources. Stored batch results share the local 1-hour result store that
read_result reads, with the same eviction, so page through a batch within the hour.
Like batch_scrape, it takes max_inline_chars (default 40,000): a page over the
limit comes back as a preview plus a result_handle for read_result.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago Changed · +2 lines 2dbb3adcacaa
- 3d ago Changed · +1 lines 4ab8cfb0d3e1
- 4d ago Changed 3cdca32a9aac
- 9d ago First seen · 127 lines · 132 tokens per session scan A bc925a7ca024
crawlforge-batch-automation is a skill published in the GitHub repository mysleekdesigns/crawlforge-mcp (2 stars, last pushed 2d ago), licensed MIT. It adds 132 tokens to every session and 1,406 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
nimble-web-expert
Get web data now — fast, incremental, immediately responsive to what the user needs. The only way Claude can access live websites. USE FOR: Fetching any URL or reading any webpage Scraping prices, listings, reviews, jobs, stats, docs from any site Running Extraction Templates — reusable, site-specific structured…
crw-dynamic-search
Programmatic web search and scrape with context isolation. Use for any research task where you need to search the web, filter results, and extract specific information — without flooding your context window with raw HTML and boilerplate. This is the single biggest token-saver in the crw skill set. Triggered by "search…
crw
Scrape, crawl, map, and search the web using fastCRW's native /v1 API. Use when the user needs web page content, site-wide extraction, URL discovery, or web search results. Single binary, 14 MB RAM; /v2 exists separately for Firecrawl migration.
crw
Scrape, crawl, map, and search the web using fastCRW's native /v1 API. Use when the user needs web page content, site-wide extraction, URL discovery, or web search results. Single binary, 14 MB RAM; /v2 exists separately for Firecrawl migration.
crw-watch
Detect what changed between two page snapshots with fastCRW — stateless diff as a REST primitive. Use when you need to track content changes, monitor a page for updates, or build a cron-based alert system: "has this page changed?", "alert me when pricing changes", "diff this week's scrape against last week's". Step 7…
chrome-devtools
Use Chrome DevTools MCP to control and inspect a live Chrome instance for network, console, performance, rendering, and Deep debugging. Pairs with playwright-cli (deterministic interaction/E2E) — complements, not duplicates.