scraper-models

scraper-models is a cursor rule for Cursor from Solihatun1/AI-Cursor-Scraping-Assistant. It costs 1,899 tokens per session, scanned A, a copy of scraper-models, MIT.

A set of coding instructions for building web scrapers in Python, including scrapers for product-list pages and product-detail pages. A product-list page shows many catalogue items, while a product-detail page describes one item.

In plain words
What is it for?
Use it when creating e-commerce scrapers that collect names, prices, links, images, categories, product details, sizes, colours, or other listed fields.
Why use it?
It gives the scraper a consistent data shape and clarifies whether it should collect catalogue entries only or also open each product page.

Cursor rule for Cursor

Written for Cursor: a Cursor rule (.mdc).

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add rules/solihatun1/ai-cursor-scraping-assistant/scraper-models
Clone the repo
git clone --depth 1 https://github.com/Solihatun1/AI-Cursor-Scraping-Assistant

Made for: Cursor.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for scraper-models

README.md
[![agentmods](https://agentmods.dev/badge/rules/solihatun1/ai-cursor-scraping-assistant/scraper-models.svg)](https://agentmods.dev/rules/solihatun1/ai-cursor-scraping-assistant/scraper-models)
Your own site
<a href="https://agentmods.dev/rules/solihatun1/ai-cursor-scraping-assistant/scraper-models"><img src="https://agentmods.dev/badge/rules/solihatun1/ai-cursor-scraping-assistant/scraper-models.svg" alt="Measured on agentmods" height="20"></a>
Per session 1,899 This file is loaded in full into every session.
When invoked 1,899 The same file — it is already loaded in full.
Security scan A 0 findings. Scan, not verified.
Origin 100% copy Near-identical to another mod in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.01899 $0.01899
Opus 5 $0.00949 $0.00949
Sonnet 5 $0.00380 $0.00380
Haiku 4.5 $0.00190 $0.00190

Measured 6d ago against content hash 0b38ed5a7e92, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-06, from the pricing page.

Security

Grade A, and why

scraper-models scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

Origin

This is a copy

100% identical to scraper-models — 0 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.

cursor-rules/scraper-models.mdc · 96 lines

How it starts

The opening of the file, as written. The whole thing — 96 lines — stays where its author put it; the contents beside it link to each section on GitHub.


description: This rule provides the description of the possible scraper types that can be created. globs: **/*.py

Scraper types

Here's a list of possible scraper types and ther data structures

E-commerce PLP

  • The structure of the items in a PLP scraper is the following: website_name, extraction_date, product_code, item_url, full_price, price, currency, image_url, brand, product_category1, product_category2, product_category3, product_name
  • When asked to create an e-commerce PLP scraper, set items.py file and in the scraper ouput accordingly to the data structure.
  • PLP pages are the product list pages in an e-commerce, also called catalogue pages. In a e-commerce PLP scraper, the scraper should crawl all the product catalog without entering the pages with product details.
  • The scraper will usually start from a home page.

E-commerce PDP

  • The strucutre of the items in a PDP scraper is the following: website_name, extraction_date, product_code, item_url, full_price, price, currency, image_url, brand, product_category1, product_category2, product_category3, product_name, product_description, product_size, product_color, additional_info
  • When asked to create an e-commerce PDP scraper, set items.py file and in the scraper ouput accordingly to the data structure.
  • PDP pages are product detail pages, the final leaf of an e-commerce website. An e-commerce PDP scraper will have in input a list of PDP pages and won't need to crawl the website further.

How to fill the scraper fields with values

  • When asked to map a field of a data structure to the information contained in the HTML, use the following rules:
    • website_name: this is a fixed value per each scraper, usually the website's name in upper case. If in doubt, ask to the operator
    • extraction_date: fixed value for the whole execution, YYYY-MM-DD format. Use datetime library
    • product_code: code that identifies every single product on the website
    • item_url: URL of the page containing the details of the product. If PDP data structure, it corresponds to response.url
    • full_price: price before the discounts. If there's no discount on the item, it's the selling price.
    • price: final selling price after the discounts. If no discount is on the website, it's the selling price.
    • currency: ISO3 Code for currency, fixed value for a whole scraper. Detect the currency from the HTML and use the ISO Code to populate the field
    • brand: brand or producer of the product sold on the website
    • product_category1: first level or product categorization, usually the first level of the breadcrumb of the page, if any.
    • product_category2: second level or product categorization, usually the second level of the breadcrumb of the page, if any.
    • product_category3: third level or product categorization, usually the third level of the breadcrumb of the page, if any.
    • product_name: name of the product as shown on the pages
  • In any case and in any field of a scraper, do not hardcode any value but always find a selector to get the correct one.
  • Always print in output every field of the scraper, even if it's empty.

Read the full file on GitHub · 96 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 96 lines · 1,899 tokens per session scan A 0b38ed5a7e92

Subscribe to this mod's changes

scraper-models is a cursor rule published in the GitHub repository Solihatun1/AI-Cursor-Scraping-Assistant (6 stars, last pushed 6mo ago), licensed MIT. It adds 1,899 tokens to every session, about $0.0095 per session on Opus 5. A static security scan graded it A with 0 findings. It is 100% identical to scraper-models, differing in 0 lines, and is treated as a copy.