pdf-inline-image-rag-mcp AGENTS.md

pdf-inline-image-rag-mcp AGENTS.md is an instructions file for Codex, OpenCode from Joncallim/pdf-inline-image-rag-mcp. It costs 551 tokens per session, scanned A, original, MIT.

Repository instructions for a PDF search server that indexes text and actual inline images in a local SQLite database.

In plain words
What is it for?
Building and searching PDF indexes, retrieving pages and images, storing captions, and validating changes to the package.
Why use it?
They define how to preserve page positions and image details while avoiding invented captions and unnecessary exposure of sensitive files.

Instructions file for CodexOpenCode

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add instructions/joncallim/pdf-inline-image-rag-mcp/agents-md
Clone the repo
git clone --depth 1 https://github.com/Joncallim/pdf-inline-image-rag-mcp

Made for: Codex, OpenCode.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for pdf-inline-image-rag-mcp AGENTS.md

README.md
[![agentmods](https://agentmods.dev/badge/instructions/joncallim/pdf-inline-image-rag-mcp/agents-md.svg)](https://agentmods.dev/instructions/joncallim/pdf-inline-image-rag-mcp/agents-md)
Your own site
<a href="https://agentmods.dev/instructions/joncallim/pdf-inline-image-rag-mcp/agents-md"><img src="https://agentmods.dev/badge/instructions/joncallim/pdf-inline-image-rag-mcp/agents-md.svg" alt="Measured on agentmods" height="20"></a>
Per session 551 This file is loaded in full into every session.
When invoked 551 The same file — it is already loaded in full.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00551 $0.00551
Opus 5 $0.00275 $0.00275
Sonnet 5 $0.00110 $0.00110
Haiku 4.5 $0.00055 $0.00055

Measured 4d ago against content hash 6846bf7c210b, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

pdf-inline-image-rag-mcp AGENTS.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

AGENTS.md · 37 lines

How it starts

The opening of the file, as written. The whole thing — 37 lines — stays where its author put it; the contents beside it link to each section on GitHub.

PDF Inline Image RAG MCP Agent Guide

This Python package provides an MCP server and CLI that builds local SQLite full-text indexes from PDF text and actual inline image blocks while retaining exact page-flow placeholders and bounding boxes.

Invariants

  • Extract actual PDF image blocks, not whole-page screenshots by default. Preserve page number, image index, dimensions, file path, and exact PDF bounding box in the stored record and text placeholder.
  • Keep extraction factual: the package does not invent OCR or captions. Save externally produced captions only through the existing caption path so full-text search stays synchronized.
  • Keep CLI and MCP behavior aligned for their shared operations: build, search, page retrieval, and database inspection. Direct image retrieval, uncaptioned- image listing, and caption persistence are MCP-only operations; do not expand the CLI to match them unless the user explicitly requests that product change.
  • Treat PDFs, extracted images, SQLite databases, captions, and output paths as potentially sensitive user data. Do not commit them, log their contents unnecessarily, or expose files outside the requested output/database scope.
  • Validate paths and replacement behavior before filesystem writes. Do not weaken explicit --replace semantics or allow a request to escape the intended output/database roots.

Repository Map

  • src/pdf_inline_image_rag_mcp/extractor.py: PDF extraction, image placement, SQLite schema/indexing, and retrieval behavior.
  • src/pdf_inline_image_rag_mcp/server.py: MCP tools and input/output contracts.
  • src/pdf_inline_image_rag_mcp/cli.py: command-line surface.
  • tests/test_extractor.py: focused extraction and persistence coverage.
  • pyproject.toml: Python, dependency, entry-point, and pytest configuration.

Focused Workflow

  • Keep one writer per implementation or test file. For substantive work, separate core extraction/database changes from a read-only MCP/CLI contract review.
  • Add a focused regression test for changes to coordinates, placeholder ordering, page selection, FTS updates, replacement, or path handling.
  • Filesystem scope, replacement/deletion, untrusted PDFs, SQL/FTS construction, MCP exposure, or sensitive output changes require an independent security review.

Read the full file on GitHub · 37 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 37 lines · 551 tokens per session scan A 6846bf7c210b

Subscribe to this mod's changes

pdf-inline-image-rag-mcp AGENTS.md is an instructions file published in the GitHub repository Joncallim/pdf-inline-image-rag-mcp (1 stars, last pushed 7d ago), licensed MIT. It adds 551 tokens to every session, about $0.0028 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other instructions, from other repositories

PixelRAG CLAUDE.md

Instructions for StarTrail-org/PixelRAG, covering pixelrag, layout and conventions.

StarTrail-org/PixelRAG · 371 tokens

pdf-rag CLAUDE.md

Claude Code instructions for abdellahiheiballa/pdf-rag, covering foundational rules, our relationship, proactiveness, designing software and test driven development (tdd).

abdellahiheiballa/pdf-rag · 2,282 tokens

invoice-builder copilot-instructions.md

Copilot instructions for piratuks/invoice-builder, a project described as: Invoice and quotation builder desktop app with PDF export, designed for small businesses and freelancers. Create, manage, and export invoices and quotes easily using a local database in an Electron-based app.

piratuks/invoice-builder · 1,320 tokens

docs-to-pdf CLAUDE.md

Claude Code instructions for jean-humann/docs-to-pdf, covering claude development guide for docs-to-pdf, project overview, development environment setup, using mise (recommended) and install mise (if not already installed).

jean-humann/docs-to-pdf · 4,714 tokens

PDF-Writer CLAUDE.md

Instructions for galkahana/PDF-Writer, covering claude code context - pdf-writer development guide, project overview, coding standards discovered, project structure and key components.

galkahana/PDF-Writer · 1,432 tokens

notebooklm-wiki-pipeline CLAUDE.md

Instructions for capitalparser/notebooklm-wiki-pipeline, covering 05notebooklmwikipipeline — 프로젝트 컨텍스트, 핵심 문제, 아키텍처, 도구 구성 and 슬래시 커맨드.

capitalparser/notebooklm-wiki-pipeline · 571 tokens