Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/ivbeg/newsworker/extractgit clone --depth 1 https://github.com/ivbeg/newsworkerWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00005 | $0.00824 |
| Opus 5 | $0.00003 | $0.00412 |
| Sonnet 5 | $0.00001 | $0.00165 |
| Haiku 4.5 | $0.00001 | $0.00082 |
Grade A, and why
extract scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
extract
Extracts news items from an HTML page and renders them in the chosen format.
newsworker extract URL [OPTIONS]
| Option | Alias | Default | Description |
|---|---|---|---|
--format |
-f |
json |
Output format: json, jsonfeed, rss, atom, csv, html, markdown, yaml. |
--output |
-o |
(stdout) | Write the result to a file instead of printing it. |
--spec |
-s |
— | Path to a YAML spec produced by analyze. |
--limit |
-n |
— | Maximum number of items to emit. |
--max-pages |
1 |
Follow up to N "next" links, merging items across pages. | |
--since |
— | Only items on or after this date (YYYY-MM-DD). |
|
--until |
— | Only items on or before this date (YYYY-MM-DD). |
|
--full-text |
false |
Follow each item link and extract the article body into content (newsworker[fulltext]). |
|
--file |
— | Local HTML file (- for stdin). Requires --base-url. |
|
--base-url |
— | Absolute HTTP(S) URL used to resolve links in local HTML. | |
--undated |
false |
Opt in to the undated listing fallback. | |
--render |
false |
Force Playwright rendering (newsworker[browser]). |
|
--explain / --explain-json |
false |
Print extraction diagnostics. | |
--user-agent |
(built-in) | Override the User-Agent used for fetching. | |
--language |
(auto) | Override the auto-detected feed language (e.g. en, fr). |
|
--proxy |
— | Proxy URL for outgoing requests. | |
--timeout |
30 |
HTTP request timeout in seconds. | |
--header |
— | Extra HTTP header Key: Value (repeatable). |
|
--cookies |
— | Path to a Netscape/Mozilla cookie jar file. | |
--insecure |
false |
Disable TLS certificate verification for this run. | |
--ignore-robots |
false |
Fetch even when robots.txt disallows it. |
|
--json-logs |
false |
Emit logs as structured JSON. | |
--no-cache |
false |
Bypass the spec and content caches for this run. | |
--refresh |
false |
Force re-fetching the page, ignoring cached content. | |
--config |
-c |
(default) | Path to a settings YAML file. |
--verbose |
-v |
false |
Verbose logging. |
Examples:
newsworker extract "https://example.com/news"
newsworker extract "https://example.com/news" -f rss
newsworker extract "https://example.com/news" -f atom -o feed.xml
newsworker extract "https://example.com/news" -s example.yaml -f rss
newsworker extract --file archive/page.html --base-url https://example.com/news/ -f json
See output formats, local input, undated listings, and diagnostics.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 55 lines · 5 tokens per session scan A 0a8436a3ff1f
extract is a command published in the GitHub repository ivbeg/newsworker (87 stars, last pushed 13d ago), licensed MIT. It adds 5 tokens to every session and 824 once invoked, about $0.0000 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
gen-research
Research a topic with primary sources and write an HTML article. Asks first whether to run locally (in-chat, 5-15 min) or route to the GitHub workflow (self-hosted runner, 20-45 min, deeper output).
gen-research-route
Route a generative-research topic to the GitHub Actions workflow (deep pipeline, 20-45 min on the self-hosted runner).
gen-research-tweet
Route a Twitter/X status URL to the generative-research GitHub workflow and expand it into a full article.
release-prepare
Require one explicit semantic version argument. Refuse an existing published version or Git tag.
release-verify
Read AGENTS.md, the entire .agents/PRIVATERELEASEPLAN.md, .agents/FAILURERETROSPECTIVE.md, and docs/releasing.md. Inspect the worktree and version-bearing files first.
setup
Install & set up newsline — rotating regional news in your status line (keeps your existing status line).