ocr skills

169 tagged ocr, measured the same way as everything else here.

Browse within: obsidian 37knowledge-management 30alchemy 20ancient-texts 20digital-humanities 20manuscripts 20nextjs 20vision 20Multimodal 19agent-skill 18GLM 17document-parsing 17claude-skill 16document-processing 16

PaddlePaddle/PaddleOCR

Skill Claude CodeCodex

Use this skill to extract structured Markdown/JSON from PDFs and document images—tables with cell-level precision, formulas as LaTeX, figures, seals, charts, headers/footers, multi-column layout and correct reading order. Trigger terms: 文档解析, 版面分析, 版面还原, 表格提取, 公式识别, 多栏排版, 扫描件结构化, 发票, 财报, 复杂 PDF, PDF转Markdown, 图表…

89k +51 1mo ago A 140 tokens original Apache-2.0

PaddlePaddle/PaddleOCR

Skill Claude CodeCodex

Use this skill whenever the user wants text extracted from images, photos, scans, screenshots, or scanned PDFs. Returns exact machine-readable strings with line-level text and optional bbox coordinates. Strong accuracy for CJK, small print, and handwritten text. Trigger terms: OCR, 文字识别, 图片转文字, 截图识字, 提取图中文字, 扫描识字, 识字…

89k +51 1mo ago A 125 tokens original Apache-2.0

Sumanth077/Hands-On-AI-Engineering

Skill Claude CodeCodex

Fetches articles from 92 Karpathy-curated RSS feeds, scores them with an LLM, selects the top 3, and delivers a formatted digest to Telegram every morning.

3.0k 7d ago A 42 tokens

verify

04

landing-ai/ade-cli

Skill Claude CodeCodex

Verify ade changes end-to-end — seed a store offline, run the CLI, drive generated artifacts in a browser.

2.4k 13d ago A 25 tokens original Apache-2.0

ade

05

landing-ai/ade-cli

Skill Claude CodeCodex

Parse documents and extract schema-shaped data with the ADE (Agentic Document Extraction) v2 APIs through the ade CLI. A local job-item store makes every run idempotent, resumable, and citable — repeat runs are free, interrupted runs resume, and every answer can cite element ids with visual evidence.

2.4k 13d ago A 65 tokens original Apache-2.0

glm-image-gen

09

zai-org/GLM-skills

Skill Claude CodeCodex

Official skill for generating high-quality images from text prompts using ZhiPu GLM-Image API. Excellent at scientific illustrations, high-quality portraits, social media graphics, and commercial posters. Supports multiple aspect ratios, HD quality, and watermark control. Use this skill when the user wants to generate…

468 4mo ago A 80 tokens original Apache-2.0

glmv-stock-analyst

10

zai-org/GLM-skills

Skill Claude CodeCodex

A stock-analysis workflow for Hong Kong, mainland Chinese, and United States shares. It combines company information, price charts, trading activity, news, and wider economic factors into a report.

468 4mo ago A 178 tokens original Apache-2.0

zai-org/GLM-skills

Skill Claude CodeCodex

Frontend visual replication skill. Explores a target website’s publicly visible pages via Playwright MCP or agent-browser, captures screenshots and layout information, then generates a static or client-side frontend replica that approximates the original’s visual appearance and page structure. This skill replicates…

468 4mo ago A 151 tokens original Apache-2.0

ddddocr

12

86maid/ddddocr

Skill Claude CodeCodex

DDDDOCR OCR recognition service with MCP protocol support. Provides optical character recognition, object detection, and slide matching capabilities. Use for: Recognizing text from captcha images, Detecting objects/text regions in images, Matching slide positions for verification codes, Performing any OCR-related…

347 4mo ago B 61 tokens

lecture-to-notes

13

drpwchen/lecture-to-notes

Skill Claude CodeCodex

Turn a lecture/conference recording (video or audio: MOV/MP4/M4A/MP3/WAV) into structured vault notes via local GPU transcription + slide extraction — 演講影片, 演講音檔, 上課錄影, '整理演講', '影片轉筆記', '音檔轉筆記', or a dropped media file. Handles batch runs.

100 25d ago A 87 tokens original MIT

figure-remap

14

drpwchen/textbook-to-note

Skill Claude CodeCodex

Extract single figures from PDF textbooks on-demand with built-in QC verification. Primary use: when writing/supplementing a note that needs an embedded figure (anatomy, classification, algorithm, imaging). Each call checks an existing fast-path crop, re-extracts from PDF if missing/wrong, retries with local-vision…

99 5d ago A 126 tokens original MIT

textbook-to-md

15

drpwchen/textbook-to-note

Skill Claude CodeCodex

Convert PDF/EPUB textbooks to searchable markdown files for an AI agent's own reference. Use this skill whenever: (1) the user asks to convert a textbook/PDF chapter to markdown, (2) you need to search textbook content and no markdown version exists yet, (3) batch-converting a set of reference books into a knowledge…

99 5d ago A 90 tokens original MIT

asdhabdua/bilibili-video-notes-skill

Skill Claude CodeCodex

A tool that turns educational or lecture videos from Bilibili, a Chinese video-sharing site, into DOCX study notes. It uses subtitles, screenshots, text recognition from images, and visual review to build the notes.

44 1mo ago A 58 tokens

vision-multimodal

17

Yts1919/dsh-vision-complete

Skill Claude CodeCodex

A skill that adds image, video, audio, document, and screenshot understanding to a text-only model. It includes tasks such as reading text from images, locating objects, transcribing speech, and analyzing media.

42 16d ago A 112 tokens original MIT

image-ppt-king

18

TateZhouSiu/image-ppt-king

Skill Claude CodeCodex

Convert slide/page images into editable PowerPoint PPTX decks by using Image Split visual layers, region schemas, transparent editable text, and QA gates. Use when Codex is given PNG/JPG/Image2/AI-generated slide images and asked to reconstruct image-based PPT pages, keep text editable, preserve layout, or turn flat…

34 3mo ago A 83 tokens original MIT

Aidenwu0209/PaddleOCR-Skills

Skill Claude CodeCodex

Install and configure two PaddleOCR Agent Skills for text recognition and structured document parsing in Codex, Claude Code, GitHub Copilot, Cursor, OpenCode, OpenClaw, and other compatible agents. Use for OCR and image-to-text from screenshots, photos, scans, and PDFs; Chinese/CJK text and bounding boxes…

34 17d ago A 102 tokens original Apache-2.0

Aidenwu0209/PaddleOCR-Skills

Skill Claude CodeCodex

Use this skill to extract structured Markdown/JSON from PDFs and document images—tables with cell-level precision, formulas as LaTeX, figures, seals, charts, headers/footers, multi-column layout and correct reading order. Trigger terms: 文档解析, 版面分析, 版面还原, 表格提取, 公式识别, 多栏排版, 扫描件结构化, 发票, 财报, 复杂 PDF, PDF转Markdown, 图表…

34 17d ago A 140 tokens original Apache-2.0

Aidenwu0209/PaddleOCR-Skills

Skill Claude CodeCodex

Use this skill whenever the user wants text extracted from images, photos, scans, screenshots, or scanned PDFs. Returns exact machine-readable strings with line-level text and optional bbox coordinates. Strong accuracy for CJK, small print, and handwritten text. Trigger terms: OCR, 文字识别, 图片转文字, 截图识字, 提取图中文字, 扫描识字, 识字…

34 17d ago A 125 tokens original Apache-2.0

frameproof

22

edvardgrishin27/frameproof

Skill Claude CodeCodex

A video analysis tool that indexes speech and visible screen text in recordings from YouTube, Loom, Zoom, Kinescope, or local MP4 files. It links claims about the screen to exact frames and time codes, while reporting sections that were not visually covered.

24 5d ago A 228 tokens original MIT

frameproof

23

edvardgrishin27/frameproof

Skill Claude CodeCodex

A video-analysis skill that indexes speech and screen content in videos, including online recordings and local files, and links findings to timestamps and frames.

24 5d ago A 228 tokens copy · 100% MIT

duepocket

24

Jiaye1998/duepocket

Skill Claude CodeCodex

Track upcoming subscription renewals, free-trial conversions, and app-wallet balances from pasted text or screenshots. Use when the user mentions a subscription, renewal, billing date, free trial ending, app wallet / stored-value balance, 续费 / 订阅 / 试用到期 / app 余额 / 充值, or wants a renewal radar, a monthly subscription…

22 2mo ago A 104 tokens