e2e-ocr

End-to-end test rules for checking text drawn on a canvas, such as PDF pages, with Tesseract.js optical character recognition.

In plain words
What is it for?
Use them in Cypress tests to wait for canvas content, convert it to an image, recognise its text, and assert the result.
Why use it?
Text painted on a canvas is not available in the page's normal HTML, so ordinary browser assertions cannot read it.

Cursor rule for Cursor

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add rules/nerds-odd-e/doughnut/e2e-ocr
Clone the repo
git clone --depth 1 https://github.com/nerds-odd-e/doughnut

Made for: Cursor.

Per session 0 Nothing until a file matches its globs; then the whole rule loads.
When invoked 383 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00000 $0.00383
Opus 5 $0.00000 $0.00192
Sonnet 5 $0.00000 $0.00077
Haiku 4.5 $0.00000 $0.00038

Measured 2d ago against content hash 1b07fc69c735, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

e2e-ocr scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.cursor/rules/e2e-ocr.mdc · 31 lines

What it actually says

E2E OCR Rules

Use this rule when asserting text that is not in the DOM, such as PDF content drawn only to <canvas>. Keep OCR in test infrastructure, not product code.

Tesseract Setup

  • Use tesseract.js, a root devDependency, for canvas-only assertions via cy.task.
  • Register a Node task in e2e_test/config/common.ts.
  • The task should call tesseract.js createWorker and recognize on image bytes, such as base64 PNG without the data:image/png;base64, prefix.
  • Example task name: ocrCanvasImage.

Language Data

  • Commit e2e_test/tesseract/eng.traineddata uncompressed.
  • Pass langPath and cachePath to that directory when creating the worker.
  • This prevents Tesseract from writing eng.traineddata to the repo root. Root-level traineddata files are gitignored as a fallback.

Page Object Pattern

  • Wait until the canvas has real ink before OCR. Sampling getImageData for dark pixels is better than checking only non-zero alpha, because an empty white fill can still have alpha.
  • Export with toDataURL("image/png").
  • Strip the data URL prefix.
  • Use cy.task(..., { timeout: ... }) for slow OCR.
  • Assert the returned string contains the expected substring with a clear message.
  • Reference: e2e_test/start/pageObjects/bookReadingPage.ts, from expectPdfBeginningVisible through expectCurrentPage(1).expectVisibleOCRContains.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 31 lines · 0 tokens per session scan A 1b07fc69c735

Subscribe to this mod's changes

e2e-ocr is a cursor rule published in the GitHub repository nerds-odd-e/doughnut (49 stars, last pushed 2d ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 383 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.