document parsing skills

21 tagged document parsing, measured the same way as everything else here.

Browse within: ocr 17data-extraction 9document-intelligence 9docx 9lts 9

PaddlePaddle/PaddleOCR

Skill Claude CodeCodex

Use this skill to extract structured Markdown/JSON from PDFs and document images—tables with cell-level precision, formulas as LaTeX, figures, seals, charts, headers/footers, multi-column layout and correct reading order. Trigger terms: 文档解析, 版面分析, 版面还原, 表格提取, 公式识别, 多栏排版, 扫描件结构化, 发票, 财报, 复杂 PDF, PDF转Markdown, 图表…

89k +128 1mo ago A 140 tokens original Apache-2.0

PaddlePaddle/PaddleOCR

Skill Claude CodeCodex

Use this skill whenever the user wants text extracted from images, photos, scans, screenshots, or scanned PDFs. Returns exact machine-readable strings with line-level text and optional bbox coordinates. Strong accuracy for CJK, small print, and handwritten text. Trigger terms: OCR, 文字识别, 图片转文字, 截图识字, 提取图中文字, 扫描识字, 识字…

89k +128 1mo ago A 125 tokens original Apache-2.0

verify

03

landing-ai/ade-cli

Skill Claude CodeCodex

Verify ade changes end-to-end — seed a store offline, run the CLI, drive generated artifacts in a browser.

2.4k +1 14d ago A 25 tokens original Apache-2.0

ade

04

landing-ai/ade-cli

Skill Claude CodeCodex

Parse documents and extract schema-shaped data with the ADE (Agentic Document Extraction) v2 APIs through the ade CLI. A local job-item store makes every run idempotent, resumable, and citable — repeat runs are free, interrupted runs resume, and every answer can cite element ids with visual evidence.

2.4k +1 14d ago A 65 tokens original Apache-2.0

Aidenwu0209/PaddleOCR-Skills

Skill Claude CodeCodex

Install and configure two PaddleOCR Agent Skills for text recognition and structured document parsing in Codex, Claude Code, GitHub Copilot, Cursor, OpenCode, OpenClaw, and other compatible agents. Use for OCR and image-to-text from screenshots, photos, scans, and PDFs; Chinese/CJK text and bounding boxes…

34 17d ago A 102 tokens original Apache-2.0

Aidenwu0209/PaddleOCR-Skills

Skill Claude CodeCodex

Use this skill to extract structured Markdown/JSON from PDFs and document images—tables with cell-level precision, formulas as LaTeX, figures, seals, charts, headers/footers, multi-column layout and correct reading order. Trigger terms: 文档解析, 版面分析, 版面还原, 表格提取, 公式识别, 多栏排版, 扫描件结构化, 发票, 财报, 复杂 PDF, PDF转Markdown, 图表…

34 17d ago A 140 tokens original Apache-2.0

Aidenwu0209/PaddleOCR-Skills

Skill Claude CodeCodex

Use this skill whenever the user wants text extracted from images, photos, scans, screenshots, or scanned PDFs. Returns exact machine-readable strings with line-level text and optional bbox coordinates. Strong accuracy for CJK, small print, and handwritten text. Trigger terms: OCR, 文字识别, 图片转文字, 截图识字, 提取图中文字, 扫描识字, 识字…

34 17d ago A 125 tokens original Apache-2.0

api-server-mcp

08

kreuzberg-dev/kreuzberg-lts

Skill Claude CodeCodex

Axum server design for document extraction endpoints, middleware, async processing, and Model Context Protocol integration for AI agents.

12 +1 19d ago A 9 tokens original MIT

kreuzberg

10

kreuzberg-dev/kreuzberg-lts

Skill Claude CodeCodex

Extract text, tables, metadata, and images from 91+ document formats (PDF, Office, images, HTML, email, archives, academic) using Kreuzberg. Use when writing code that calls Kreuzberg APIs in Python, Node.js/TypeScript, Rust, or CLI. Covers installation, extraction (sync/async), configuration (OCR, chunking, output…

12 +1 19d ago A 89 tokens original MIT

legal-kb-builder

11

Youchu-lawhub/legal-kb-builder

Skill Claude CodeCodex

A tool for turning legal documents and other source material into a searchable local knowledge base. It supports sources such as PDF, DOCX, images, Markdown, and question-and-answer sets.

6 1mo ago B 124 tokens

legalgraph

12

gauravmanandhar/legal-graphify-skill

Skill Claude CodeCodex

Turn legal documents, policies, contracts, compliance rules, and regulatory text into a queryable knowledge graph. Extracts clauses, obligations, prohibitions, permissions, cross-references, and detects conflicts. Answers questions like 'What does policy X say about Y?', 'Find all obligations for role Z', 'Are these…

5 1mo ago A 70 tokens original Apache-2.0

paddleocr-mcp

13

Nicvank/paddleocr-mcp

Skill Claude CodeCodex

Use the local PaddleOCR MCP SDK v2 server for image OCR, PDF/document parsing, screenshots, tables, and structured text extraction.

2 1mo ago A 33 tokens original MIT

vision-augment

14

CaoMeiYouRen/vision-augment

Skill Claude CodeCodex

A visual and document-analysis service for agents that cannot directly understand images. It can inspect pictures, read text with OCR, and parse documents such as PDF, Word, PowerPoint, spreadsheet, and HTML files.

0 21d ago A 104 tokens original MIT

dbv-pdf2md

15

davidbuenov/dbv-pdf2md

Skill Claude CodeCodex

Conversor avanzado de PDF a Markdown con extracción física de imágenes, detección y reconstrucción de hipervínculos ("pincha aquí") e integración MCP.

0 25d ago A 39 tokens original MIT