Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add xberg-io/crawlberg --skill automating-the-browsergit clone --depth 1 https://github.com/xberg-io/crawlbergWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/xberg-io/crawlberg/automating-the-browser)<a href="https://agentmods.dev/skills/xberg-io/crawlberg/automating-the-browser"><img src="https://agentmods.dev/badge/skills/xberg-io/crawlberg/automating-the-browser/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/xberg-io/crawlberg/automating-the-browser"><img src="https://agentmods.dev/badge/skills/xberg-io/crawlberg/automating-the-browser.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00063 | $0.01377 |
| Opus 5 | $0.00032 | $0.00688 |
| Sonnet 5 | $0.00013 | $0.00275 |
| Haiku 4.5 | $0.00006 | $0.00138 |
Grade A, and why
automating-the-browser scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 138 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Automating the browser
crawlberg interact <url> --actions '[...]' drives a real headless browser
through an ordered list of actions, then captures the resulting page. Reach for
it when a static scrape is not enough: the content lives behind a "Load more"
button, an infinite scroll, a form submission, or JS that only runs on
interaction.
Quick recipe
crawlberg interact https://example.com \
--actions '[{"type":"click","selector":"#load-more"},
{"type":"wait","milliseconds":500},
{"type":"scrape"}]'
Actions run in order. The result wraps the final page state under
interaction (see Output below).
Flag surface
| Flag | Default | Purpose |
|---|---|---|
--actions |
— | Required. JSON array of action objects (see below). |
--format |
json |
json (full result) or markdown (final HTML only). |
--timeout |
30000 |
Per-request timeout in ms. |
--browser-mode |
auto |
auto, always, never. Interaction needs a browser. |
--browser-endpoint |
— | External CDP ws:// or wss:// URL. |
--config |
— | Inline JSON or @file.json for the full CrawlConfig. |
There is no --respect-robots-txt flag on interact; it targets the one URL
you point it at.
Action schema
Each action is a JSON object tagged by type (camelCase). Unknown fields are
rejected, so match the shapes exactly:
type |
Fields | Notes |
|---|---|---|
click |
selector |
Click the element matching the CSS selector. |
type |
selector, text |
Type text into the input matching selector. |
press |
key |
Press a key, e.g. "Enter", "Tab", "Escape". |
scroll |
direction ("up"/"down"), selector?, amount? |
Scroll the page, or a scrollable element if selector is set. |
wait |
milliseconds?, selector? |
Wait a fixed time, or until selector appears (selector wins). |
screenshot |
fullPage? |
Capture the viewport, or the full scrollable page if true. |
executeJs |
script |
Run arbitrary JS in the page context. Trusted scripts only. |
scrape |
— | Capture the current page HTML into the result. |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 138 lines · 63 tokens per session scan A ae1804941aa8
automating-the-browser is a skill published in the GitHub repository xberg-io/crawlberg (169 stars, last pushed 3d ago), licensed MIT. It adds 63 tokens to every session and 1,377 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
tree-sitter-language-pack
Parse and extract code intelligence from 371 programming languages using tree-sitter grammars. Use when writing code that parses source, extracts structure/imports/exports/symbols/docstrings/comments, detects a language, runs syntax diagnostics, or produces syntax-aware chunks for LLMs — in Rust, Python…
extracting-code-structure
Use when the user wants structured code metadata from a source file — functions, classes, imports, exports, symbols, docstrings, comments, or syntax diagnostics. Covers ts-pack process feature flags, the JSON result shape, and the default feature set.
chunking-for-llms
Use when the user wants to split source code into chunks for an LLM context window without breaking syntax mid-construct. Covers ts-pack process --chunk-size, why syntax-aware splits beat fixed-byte splits, picking a size, and the chunk JSON shape.
detecting-languages
Use when the user wants to know which programming language a file or snippet is. Covers implicit detection in ts-pack parse/process, confirming support with ts-pack list/info, and the SDK detection functions for path, extension, and raw content.
managing-parsers
Use when the user needs to manage the tree-sitter parser cache — prefetch parsers for offline or CI runs, list what is downloaded, inspect a language, find the cache directory, or clean it. Covers ts-pack download, list, info, cache-dir, clean, and init.
parsing-source
Use when the user wants a tree-sitter syntax tree for a source file — an s-expression dump or JSON tree. Covers ts-pack parse, language auto-detection vs --language, stdin input, and reading haserrors.