text extraction skills

12 tagged text extraction, measured the same way as everything else here.

Browse within: RAG 10hocr 10html-converter 10markdown-converter 10text-processing 10

converting-html

01

xberg-io/html-to-markdown

Skill Claude CodeCodex

Use when converting HTML to Markdown, Djot, or plain text. Covers output formats, heading and code-block styles, lists, escaping, wrapping, and HTML preprocessing.

860 3d ago A 38 tokens original MIT

extracting-metadata

02

xberg-io/html-to-markdown

Skill Claude CodeCodex

Use when extracting metadata from HTML — title, description, language, Open Graph, JSON-LD / Microdata / RDFa, headers, links, and images. Covers the --json output shape and the --extract-metadata flag.

860 3d ago A 52 tokens original MIT

html-to-markdown

03

xberg-io/html-to-markdown

Skill Claude CodeCodex

Convert HTML to Markdown, Djot, or plain text with structured extraction. Use when writing code that calls html-to-markdown APIs in Rust, Python, TypeScript, Go, Ruby, PHP, Java, C#, Elixir, R, C, or WASM. Covers installation, conversion, configuration, metadata extraction, tables, document structure, inline images…

860 3d ago A 85 tokens original MIT

pdf-ocr-skill

04

yejinlei/pdf-ocr-skill

Skill Claude CodeCodex

A PDF and image text-recognition tool for scanned documents. OCR, or optical character recognition, converts text visible in scans or pictures into machine-readable text, including Chinese and English.

14 4mo ago A 58 tokens original MIT