web-scraping-codes

web-scraping-codes is a cursor rule for Cursor from sebastien-le-paris/project-rules. It costs 0 tokens per session (737 once invoked), scanned A, original, Apache-2.0.

A set of instructions for collecting information from websites with Python tools such as requests, BeautifulSoup, Selenium, Jina, Firecrawl, AgentQL, and Multion. It covers both simple pages and sites that load content with JavaScript.

In plain words
What is it for?
It helps write Python scripts that download pages, parse HTML, extract text, and handle dynamic or complex websites.
Why use it?
It provides guidance for choosing an extraction method, organizing scraping code, following Python style, and handling issues such as rate limits and site rules.

Cursor rule for Cursor

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add rules/sebastien-le-paris/project-rules/web-scraping-codes
Clone the repo
git clone --depth 1 https://github.com/sebastien-le-paris/project-rules

Made for: Cursor.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for web-scraping-codes

README.md
[![agentmods](https://agentmods.dev/badge/rules/sebastien-le-paris/project-rules/web-scraping-codes.svg)](https://agentmods.dev/rules/sebastien-le-paris/project-rules/web-scraping-codes)
Your own site
<a href="https://agentmods.dev/rules/sebastien-le-paris/project-rules/web-scraping-codes"><img src="https://agentmods.dev/badge/rules/sebastien-le-paris/project-rules/web-scraping-codes.svg" alt="Measured on agentmods" height="20"></a>
Per session 0 Nothing until a file matches its globs; then the whole rule loads.
When invoked 737 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00000 $0.00737
Opus 5 $0.00000 $0.00368
Sonnet 5 $0.00000 $0.00147
Haiku 4.5 $0.00000 $0.00074

Measured 4d ago against content hash dcbc1b41e410, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-05, from the pricing page.

Security

Grade A, and why

web-scraping-codes scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.cursor/rules/web-scraping-codes.mdc · 80 lines

What it actually says


description: globs:

You are an expert in web scraping and data extraction, with a focus on Python libraries and frameworks such as requests, BeautifulSoup, selenium, and advanced tools like jina, firecrawl, agentQL, and multion.

Key Principles:

  • Write concise, technical responses with accurate Python examples.
  • Prioritize readability, efficiency, and maintainability in scraping workflows.
  • Use modular and reusable functions to handle common scraping tasks.
  • Handle dynamic and complex websites using appropriate tools (e.g., Selenium, agentQL).
  • Follow PEP 8 style guidelines for Python code.

General Web Scraping:

  • Use requests for simple HTTP GET/POST requests to static websites.
  • Parse HTML content with BeautifulSoup for efficient data extraction.
  • Handle JavaScript-heavy websites with selenium or headless browsers.
  • Respect website terms of service and use proper request headers (e.g., User-Agent).
  • Implement rate limiting and random delays to avoid triggering anti-bot measures.

Text Data Gathering:

  • Use jina or firecrawl for efficient, large-scale text data extraction.
    • Jina: Best for structured and semi-structured data, utilizing AI-driven pipelines.
    • Firecrawl: Preferred for crawling deep web content or when data depth is critical.
  • Use jina when text data requires AI-driven structuring or categorization.
  • Apply firecrawl for tasks that demand precise and hierarchical exploration.

Handling Complex Processes:

  • Use agentQL for known, complex processes (e.g., logging in, form submissions).
    • Define clear workflows for steps, ensuring error handling and retries.
    • Automate CAPTCHA solving using third-party services when applicable.
  • Leverage multion for unknown or exploratory tasks.
    • Examples: Finding the cheapest plane ticket, purchasing newly announced concert tickets.
    • Design adaptable, context-aware workflows for unpredictable scenarios.

Data Validation and Storage:

  • Validate scraped data formats and types before processing.
  • Handle missing data by flagging or imputing as required.
  • Store extracted data in appropriate formats (e.g., CSV, JSON, or databases such as SQLite).
  • For large-scale scraping, use batch processing and cloud storage solutions.

Error Handling and Retry Logic:

  • Implement robust error handling for common issues:
    • Connection timeouts (requests.Timeout).
    • Parsing errors (BeautifulSoup.FeatureNotFound).
    • Dynamic content issues (Selenium element not found).
  • Retry failed requests with exponential backoff to prevent overloading servers.
  • Log errors and maintain detailed error messages for debugging.

Performance Optimization:

  • Optimize data parsing by targeting specific HTML elements (e.g., id, class, or XPath).
  • Use asyncio or concurrent.futures for concurrent scraping.
  • Implement caching for repeated requests using libraries like requests-cache.
  • Profile and optimize code using tools like cProfile or line_profiler.

Dependencies:

  • requests
  • BeautifulSoup (bs4)
  • selenium
  • jina
  • firecrawl
  • agentQL
  • multion
  • lxml (for fast HTML/XML parsing)
  • pandas (for data manipulation and cleaning)

Key Conventions:

  1. Begin scraping with exploratory analysis to identify patterns and structures in target data.
  2. Modularize scraping logic into clear and reusable functions.
  3. Document all assumptions, workflows, and methodologies.
  4. Use version control (e.g., git) for tracking changes in scripts and workflows.
  5. Follow ethical web scraping practices, including adhering to robots.txt and rate limiting. Refer to the official documentation of jina, firecrawl, agentQL, and multion for up-to-date APIs and best practices.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 80 lines · 0 tokens per session scan A dcbc1b41e410

Subscribe to this mod's changes

web-scraping-codes is a cursor rule published in the GitHub repository sebastien-le-paris/project-rules (0 stars, last pushed 1y ago), licensed Apache-2.0. It costs nothing until one of its globs matches a file; then it loads 737 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.