ocr skills

150 tagged ocr, measured the same way as everything else here.

Browse within: obsidian 37knowledge-management 30digital-humanities 21vision 21alchemy 20ancient-texts 20manuscripts 20nextjs 20Multimodal 19agent-skill 18GLM 17document-parsing 17claude-skill 16document-processing 16

vault-scout

25

mm-weber/loremaester

Skill Claude CodeCodex

Use when you need to find, list, or read-and-summarize Obsidian vault notes mechanically — e.g., "find all NPC notes tied to faction X", "list all session N scene files", "read frontmatter of these 12 notes and return a table", "grep all Locations/ notes for wikilinks to [[NPC Name]]". Dispatches a Haiku subagent to…

not rated 19 1mo ago A 147 tokens original MIT

worldbuilding

26

mm-weber/loremaester

Skill Claude CodeCodex

Use when creating or expanding worldbuilding content — locations, NPCs, factions, lore entries, quests, or items for a TTRPG campaign managed in an Obsidian vault. Triggers on any request to create, describe, detail, or flesh out campaign world elements, even if the user doesn't explicitly say "worldbuilding". Also…

not rated 19 1mo ago A 93 tokens original MIT

batch-translate

27

Embassy-of-the-Free-Mind/sourcelibrary-v2

Skill Claude CodeCodex

Batch process books through the complete pipeline - generate cropped images for split pages, OCR all pages, then translate with context. Use when asked to process, OCR, translate, or batch process one or more books.

not rated 17 today A 46 tokens AGPL-3.0

progress

28

Embassy-of-the-Free-Mind/sourcelibrary-v2

Skill Claude CodeCodex

Check pipeline processing progress — all phases including OCR, translation, image extraction, enrichment, and more. Shows real page-level verification, not just job counters. Use when asked "how's it going?", "progress?", "status?", "now?", or any progress check.

not rated 17 today A 56 tokens AGPL-3.0

qa-audit

29

Embassy-of-the-Free-Mind/sourcelibrary-v2

Skill Claude CodeCodex

Quality auditor for Source Library. Prioritizes verifying original language texts (not modern translations), auditing metadata accuracy against title pages, USTC alignment, and translation quality. Use for systematic quality control or to identify modern translations that should be replaced with originals.

not rated 17 today A 55 tokens AGPL-3.0

pdf-ocr-skill

30

yejinlei/pdf-ocr-skill

Skill Claude CodeCodex

A PDF and image text-recognition tool for scanned documents. OCR, or optical character recognition, converts text visible in scans or pictures into machine-readable text, including Chinese and English.

not rated 14 4mo ago A 58 tokens original MIT

api-server-mcp

31

kreuzberg-dev/kreuzberg-lts

Skill Claude CodeCodex

Axum server design for document extraction endpoints, middleware, async processing, and Model Context Protocol integration for AI agents.

not rated 12 +1 21d ago A 9 tokens original MIT

kreuzberg

33

kreuzberg-dev/kreuzberg-lts

Skill Claude CodeCodex

Extract text, tables, metadata, and images from 91+ document formats (PDF, Office, images, HTML, email, archives, academic) using Kreuzberg. Use when writing code that calls Kreuzberg APIs in Python, Node.js/TypeScript, Rust, or CLI. Covers installation, extraction (sync/async), configuration (OCR, chunking, output…

not rated 12 +1 21d ago A 89 tokens original MIT

mineru-pdf

34

hawkongz/mineru-pdf

Skill Claude CodeCodex

High-accuracy PDF content extraction using MinerU (Shanghai AI Lab). Use this whenever the user needs to extract text, formulas, tables, or images from a complex PDF — especially academic papers, multi-column layouts, scanned documents, or any PDF where pypdf produces garbled /Cxx formula output. Trigger on: "MinerU"…

not rated 10 3mo ago A 154 tokens original MIT

mineru-pdf

35

hawkongz/mineru-pdf

Skill Claude CodeCodex

Instructions for extracting text, formulas, tables, and images from difficult PDF files with MinerU, a document-reading tool. It is intended for papers, multi-column layouts, and scanned documents where basic PDF readers may produce poor results.

not rated 10 3mo ago A 180 tokens original MIT

runyuan-wang/book-to-skill-distillation

Skill Claude CodeCodex

End-to-end workflow for rewriting a book, long PDF, EPUB, manual, course, guideline library, same-domain multi-book source pack, large database, or methodology into an agent-native LingTai / Agent Skill structure. Use when the task is not a summary, but a reusable skill/knowledge system: source triage…

not rated 9 2mo ago A 120 tokens

DeHor-Labs/transcreve-ai

Skill Claude CodeCodex

Use when an AI agent invokes TranscreveAI as a nested video-to-knowledge capability and must return a durable handoff to the caller.

not rated 5 3d ago C 38 tokens original MIT

DeHor-Labs/transcreve-ai

Skill Claude CodeCodex

Use to turn video URLs or media files into evidence-backed knowledge dossiers with TranscreveAI. Trigger on Reels, YouTube, TikTok, Loom, Vimeo, X/Twitter video links, local media files, video summaries, dossier requests, RAG over video runs, or requests to use TranscreveAI.

not rated 5 3d ago C 73 tokens original MIT

pdf-table-to-excel

39

kujiangmudao/tablepack

Skill Claude CodeCodex

A workflow for turning tables in PDF files into Excel workbooks, with one sheet per table and supporting images and quality notes. It uses MinerU, a tool that extracts content from PDFs, and needs an image-reading model to check the results.

not rated 5 21d ago A 132 tokens original MIT

Aidenwu0209/dsh-PaddleOCR-Skills

Skill Claude CodeCodex

Install and configure the native PaddleOCR plugin for DeepSeek Harness (DSH) from the Settings → PaddleOCR GUI. Use for OCR and image-to-text from screenshots, scans, and PDFs; Chinese/CJK text; PDF-to-Markdown; structured document parsing with tables, formulas, layout, and reading order; or DSH endpoint, credential…

not rated 4 19d ago A 88 tokens original Apache-2.0

ocrCN

43

Agents365-ai/ocrCN

Skill Claude CodeCodex

Multi-platform Chinese OCR text recognition via PaddleOCR/Baidu/Tencent/Alibaba/EasyOCR — 5 backends, all work in China.

not rated 2 1mo ago A 31 tokens

beautiful-article

44

Linearl/reasonix_skill_repo

Skill Claude CodeCodex

A workflow for turning supplied material—such as a web page, PDF, document, Markdown, text, screenshot, or pasted content—into a single offline HTML article.

not rated 2 6d ago A 226 tokens

book-to-skill

45

Linearl/reasonix_skill_repo

Skill Claude CodeCodex

Converts books and documents (PDF, EPUB, DOCX, HTML, Markdown, plain text, RTF, MOBI/AZW with Calibre) into structured agent skills, extracting frameworks, mental models, principles, techniques, and anti-patterns. Use when the user wants to study a document through GitHub Copilot CLI, Amp, or Claude Code, apply an…

not rated 2 6d ago A 95 tokens

invest_analysis

46

Linearl/reasonix_skill_repo

Skill Claude CodeCodex

A structured research workflow for studying companies and markets in mainland China and Hong Kong. It covers industry selection, supply-chain research, financial-report checks, and investor sentiment.

not rated 2 6d ago A 75 tokens

vision

47

DDDFXYqiming/dsh-vision-skill

Skill Claude CodeCodex

A local-image analysis skill that identifies what is shown in pictures, screenshots, and error images.

not rated 2 4d ago A 50 tokens copy · 100% MIT

liusencomic-cyber/douyin-agent-kit

Skill Claude CodeCodex

A batch workflow for archiving Douyin favorites or all posts from selected creators into Obsidian. It uses account cookies, rate limits, downloading, transcription, deduplication, and optional scheduled synchronization.

not rated 2 4d ago A 35 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: