Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add rules/solihatun1/ai-cursor-scraping-assistant/scraper-modelsgit clone --depth 1 https://github.com/Solihatun1/AI-Cursor-Scraping-AssistantWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/rules/solihatun1/ai-cursor-scraping-assistant/scraper-models)<a href="https://agentmods.dev/rules/solihatun1/ai-cursor-scraping-assistant/scraper-models"><img src="https://agentmods.dev/badge/rules/solihatun1/ai-cursor-scraping-assistant/scraper-models.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.01899 | $0.01899 |
| Opus 5 | $0.00949 | $0.00949 |
| Sonnet 5 | $0.00380 | $0.00380 |
| Haiku 4.5 | $0.00190 | $0.00190 |
Grade A, and why
scraper-models scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
This is a copy
100% identical to scraper-models — 0 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
How it starts
The opening of the file, as written. The whole thing — 96 lines — stays where its author put it; the contents beside it link to each section on GitHub.
description: This rule provides the description of the possible scraper types that can be created. globs: **/*.py
Scraper types
Here's a list of possible scraper types and ther data structures
E-commerce PLP
- The structure of the items in a PLP scraper is the following: website_name, extraction_date, product_code, item_url, full_price, price, currency, image_url, brand, product_category1, product_category2, product_category3, product_name
- When asked to create an e-commerce PLP scraper, set items.py file and in the scraper ouput accordingly to the data structure.
- PLP pages are the product list pages in an e-commerce, also called catalogue pages. In a e-commerce PLP scraper, the scraper should crawl all the product catalog without entering the pages with product details.
- The scraper will usually start from a home page.
E-commerce PDP
- The strucutre of the items in a PDP scraper is the following: website_name, extraction_date, product_code, item_url, full_price, price, currency, image_url, brand, product_category1, product_category2, product_category3, product_name, product_description, product_size, product_color, additional_info
- When asked to create an e-commerce PDP scraper, set items.py file and in the scraper ouput accordingly to the data structure.
- PDP pages are product detail pages, the final leaf of an e-commerce website. An e-commerce PDP scraper will have in input a list of PDP pages and won't need to crawl the website further.
How to fill the scraper fields with values
- When asked to map a field of a data structure to the information contained in the HTML, use the following rules:
- website_name: this is a fixed value per each scraper, usually the website's name in upper case. If in doubt, ask to the operator
- extraction_date: fixed value for the whole execution, YYYY-MM-DD format. Use datetime library
- product_code: code that identifies every single product on the website
- item_url: URL of the page containing the details of the product. If PDP data structure, it corresponds to response.url
- full_price: price before the discounts. If there's no discount on the item, it's the selling price.
- price: final selling price after the discounts. If no discount is on the website, it's the selling price.
- currency: ISO3 Code for currency, fixed value for a whole scraper. Detect the currency from the HTML and use the ISO Code to populate the field
- brand: brand or producer of the product sold on the website
- product_category1: first level or product categorization, usually the first level of the breadcrumb of the page, if any.
- product_category2: second level or product categorization, usually the second level of the breadcrumb of the page, if any.
- product_category3: third level or product categorization, usually the third level of the breadcrumb of the page, if any.
- product_name: name of the product as shown on the pages
- In any case and in any field of a scraper, do not hardcode any value but always find a selector to get the correct one.
- Always print in output every field of the scraper, even if it's empty.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 96 lines · 1,899 tokens per session scan A 0b38ed5a7e92
scraper-models is a cursor rule published in the GitHub repository Solihatun1/AI-Cursor-Scraping-Assistant (6 stars, last pushed 6mo ago), licensed MIT. It adds 1,899 tokens to every session, about $0.0095 per session on Opus 5. A static security scan graded it A with 0 findings. It is 100% identical to scraper-models, differing in 0 lines, and is treated as a copy.
Other cursor rules, from other repositories
scraper-models
Cursor rule "scraper-models" from TheWebScrapingClub/AI-Cursor-Scraping-Assistant, covering scraper types, e-commerce plp, e-commerce pdp, how to fill the scraper fields with values and how to create an e-commerce plp scraper.
website-analysis
description: This rule provides a step by step guide to analyze a website and its code, in order to write a better Scrapy scraper. globs: /.py.
scrapy-step-by-step-process
description: This rule provides a step by step guide to follow for a successful Scrapy project. Read and implement the rules oncained in this file in first place. globs: /.py.
scrapy
Cursor rule "scrapy" from TheWebScrapingClub/AI-Cursor-Scraping-Assistant, covering scrapy best practices, 1. code organization and structure, 1.1. directory structure, 1.2. file naming conventions and 1.3. module organization.
prerequisites
description: This rule provides the action that should be taken before starting implementing a Scrapy spider. globs: /.py.
playwright
Playwright: e2e testing, page objects, fixtures, assertions.