Azure/gpt-rag-ingestion

The GPT-RAG Data Ingestion service automates processing of diverse documents—PDFs, images, spreadsheets, transcripts, and SharePoint—readying them for Azure AI Search. It applies smart chunking, generates text and image embeddings, and enables rich, multimodal retrieval.

About the project

GPT-RAG Data Ingestion is a service that processes documents such as PDFs, images, spreadsheets, transcripts, and SharePoint files so they can be searched through Azure AI Search. It prepares data with format-specific chunking and text or image embeddings for multimodal retrieval in agent-based applications.

These files are Azure/gpt-rag-ingestion's own configuration. They tell Claude Code, GitHub Copilot, Codex and OpenCode how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

189Stars on the repository
19Files it configures its agents with
3,887Tokens loaded in every session
4Agents configured

Instructions

Skills

Agents