ocr skills

150 tagged ocr, measured the same way as everything else here.

Browse within: obsidian 37knowledge-management 30digital-humanities 21vision 21alchemy 20ancient-texts 20manuscripts 20nextjs 20Multimodal 19agent-skill 18GLM 17document-parsing 17claude-skill 16document-processing 16

liusencomic-cyber/douyin-agent-kit

Skill Claude CodeCodex

A workflow for turning Douyin videos or image posts into text, summaries, intent judgments, and archived notes. Douyin is a Chinese social-media platform for short videos and image posts.

not rated 2 4d ago C 68 tokens original MIT

pdf-ocr-to-markdown

50

jseook11/codex-pdf-ocr-to-markdown-skill

Skill Claude CodeCodex

OCR PDFs/images into user-facing Markdown with internal JSON/quality artifacts. Preserve originals; never create or modify PDFs. Use bundled scripts, not ad hoc OCR code. Mark visual pages pending unless images are actually inspected.

not rated 2 4mo ago A 50 tokens original MIT

paddleocr-mcp

51

Nicvank/paddleocr-mcp

Skill Claude CodeCodex

Use the local PaddleOCR MCP SDK v2 server for image OCR, PDF/document parsing, screenshots, tables, and structured text extraction.

not rated 2 1mo ago A 33 tokens original MIT

vision-perceive

52

Yuhang-uestc/deepvision-local-mcp

Skill Claude CodeCodex

A multi-step local process for understanding images with a text-only AI model. It chooses a quick or detailed path and can combine image description, text reading, object finding, cropping, and checking results.

not rated 2 23d ago A 133 tokens original MIT

image-reading

53

jing-hy/picturereader-zcode

Skill Claude CodeCodex

Read and understand images like a multimodal model using the picturereader tools (imagescan / imageocr / imagesample). Applies a verified 5-step workflow (global tone → find subjects → verify text → judge material → synthesize) guided by grounded principles and cross-image insights. Use whenever you need to look at an…

not rated 2 18d ago A 72 tokens original MIT archived

omd-local/markdown-everything

Skill Claude CodeCodex

Anti-slop frontend skill for landing pages, portfolios, and redesigns. The agent reads the brief, infers the right design direction, and ships interfaces that do not look templated. Real design systems when applicable, audit-first on redesigns, strict pre-flight check.

not rated 1 19d ago A 61 tokens copy · 100% MIT

KOLRNCH/Universal-AI-Converter

Skill Claude CodeCodex

Automatically use Universal AI Converter whenever the user attaches, references, or asks to inspect a local document, spreadsheet, presentation, archive, image, scan, audio recording, or video. Trigger even when the user does not name the converter or ask for conversion explicitly.

not rated 1 1mo ago A 57 tokens original MIT

eagleeye

56

baimaomaomao556/eagleeye-mcp

Skill Claude CodeCodex

Deterministic visual toolbox via EagleEye MCP (not a CLI). Use when the task needs a live screenshot, window capture, waiting for UI to appear, pixel/color/template measurement, OCR, UI checks, or visual regression. Do not use to turn a conversation image into a one-shot description (that is modlens), or to rebuild…

not rated 0 19d ago A 117 tokens original MIT

eagleeye

57

baimaomaomao556/eagleeye-mcp

Skill Claude CodeCodex

Deterministic visual toolbox via EagleEye MCP (not a CLI). Use when the task needs a live screenshot, window capture, waiting for UI to appear, pixel/color/template measurement, OCR, UI checks, or visual regression. Do not use to turn a conversation image into a one-shot description (that is modlens), or to rebuild…

not rated 0 19d ago A 117 tokens copy · 86% MIT

manga-translate

58

chiraitori/manga-trans

Skill Claude CodeCodex

Automated scanlation and manga translation workflow using Manga Trans MCP Server. Guides chapter analysis, context-aware multilingual localization (English, Vietnamese, etc.), region-locked manifest creation, typesetting, strict pixel QA validation, and interactive Web Studio review hand-off.

not rated 0 11d ago A 56 tokens