apache/tika

The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).

About the project

Apache Tika is a toolkit that reads many kinds of files and extracts their text and descriptive metadata, including from formats such as PDF, PowerPoint, and Excel. Applications and agent pipelines use it to turn documents into content they can process, search, or pass to language models. The catalogue skills provide reusable ways for coding agents to run Tika for file-to-Markdown extraction.

These files are apache/tika's own configuration. They tell Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

4,036Stars on the repository
1Files it configures its agents with
936Tokens loaded in every session
2Agents configured

Instructions