extraction skills

33 tagged extraction, measured the same way as everything else here.

Browse within: content 12java 12metadata 12tika 12analysis 5pdf 5

oss-fuzz

01

apache/tika

Skill Claude CodeCodex

Run Tika's OSS-Fuzz Jazzer targets locally against a working-tree checkout — build the image, build fuzzers from local source, fuzz a target, run a corpus as a regression pass, reproduce a crash, and add seeds. Use for "fuzz the OneNote parser", "run OneNoteParserFuzzer against these files", "reproduce an OSS-Fuzz…

4.0k today A 90 tokens original Apache-2.0

apache/tika

Skill Claude CodeCodex

Update/publish the Apache Tika website (tika-site SVN repo) for a release — step 17 of the Release Process. Handles the 4.x track (Changes page + aggregate javadoc + Antora docs branch) vs the 3.x maintenance track (full per-version apt docs + javadoc). Use for "update the site", "publish the site for X.Y.Z", "the…

4.0k today A 92 tokens original Apache-2.0

file-to-markdown

03

apache/tika

Skill Claude CodeCodex

Turn almost any file into Markdown plus metadata — PDF, Office, HTML, email, archives, images, audio/video, 1000+ formats — powered by Apache Tika, either via the tika-app CLI (zero setup, one file) or a running tika-server (curl, warm process, many calls). Leads with rmeta (structured, embedded-item-aware output) as…

4.0k today A 147 tokens original Apache-2.0

live-neon/persona-mcp

Skill Claude CodeCodex

Automatically discover what your AI agent believes by analyzing its real outputs — Pattern-Based Distillation for agent behavior.

2 4mo ago A 27 tokens original MIT