Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add instructions/hwalde/vision-mcp/agents-mdgit clone --depth 1 https://github.com/hwalde/vision-mcpWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.01850 | $0.01850 |
| Opus 5 | $0.00925 | $0.00925 |
| Sonnet 5 | $0.00370 | $0.00370 |
| Haiku 4.5 | $0.00185 | $0.00185 |
Grade A, and why
vision-mcp AGENTS.md scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 159 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Vision MCP — Setup & Konfiguration
Diese Datei ist die einzige Quelle der Projekt-Doku.
CLAUDE.mdimportiert sie per@AGENTS.md, damit Claude Code sie automatisch mitlädt — nicht duplizieren, hier pflegen.
MCP-Server (stdio, Python/FastMCP) für Bild-Analyse: Objekterkennung (YOLOv8) und Texterkennung/OCR (EasyOCR), plus abgeleitete Layout-Aussagen (Abstände, Zentrierung). Läuft auf Windows, macOS und Linux, komplett lokal und offline — es geht kein Bild an eine Cloud-API.
Die Ausgabe ist bewusst LLM-optimierter Markdown-Text, kein JSON: natürliche Sätze wie „Text „Headline" liegt 276px oberhalb von person." verarbeitet ein LLM zuverlässiger als verschachtelte Koordinaten-Strukturen.
Was der Server braucht
| Baustein | Zweck |
|---|---|
| Python ≥ 3.10 | Laufzeit |
Pakete aus requirements.txt |
u. a. ultralytics (YOLO) und easyocr — beide ziehen torch transitiv mit |
YOLOv8-Gewichte (yolov8n.pt, ~6 MB) |
werden beim ersten Start automatisch geladen |
| EasyOCR-Sprachmodelle (~100 MB) | werden beim ersten OCR-Aufruf automatisch geladen |
Kein API-Key, kein Cloud-Konto, keine Auth. Nur der erste Start braucht Internet, um die Modelle zu holen.
Installation
Nach dem Auschecken genügen zwei Befehle — identisch auf allen Betriebssystemen:
macOS / Linux
pip3 install -r requirements.txt
python3 server.py --selftest
Windows (PowerShell)
pip install -r requirements.txt
python server.py --selftest
--selftest ist der schnellste Weg zur Gewissheit: Er prüft die Python-Version, meldet jedes
fehlende Paket namentlich, lädt YOLO- und OCR-Modelle vor (der einzige Schritt, der Internet
braucht) und analysiert ein selbst erzeugtes Testbild. Endet er mit „Alles bereit", ist der Server
einsatzfähig — man muss den Fehler nicht erst beim ersten Tool-Aufruf im Agenten entdecken.
Exit-Code 0 = alles in Ordnung, 1 = Problem.
Was dabei zu erwarten ist: ultralytics und easyocr ziehen PyTorch mit — mehrere hundert
MB, die Installation dauert entsprechend. Der erste --selftest lädt zusätzlich ~6 MB
YOLO-Gewichte und ~100 MB OCR-Modelle. Danach läuft alles offline.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 159 lines · 1,850 tokens per session scan A f62523745f0f
vision-mcp AGENTS.md is an instructions file published in the GitHub repository hwalde/vision-mcp (0 stars, last pushed 1mo ago), licensed MIT. It adds 1,850 tokens to every session, about $0.0093 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other instructions, from other repositories
vscode buildNext.instructions.md
Working notes and architecture documentation for the new esbuild-based build system in build/next. Use when making changes to the new build pipeline (transpile/bundle commands, NLS plugin, source-map handling, resource copying, or self-hosting watch tasks).
spec-kit AGENTS.md
AGENTS.md instructions for github/spec-kit, covering agents.md, about spec kit and specify, quickstart — add a new integration in 5 steps, integration architecture and integrationmanifest — file tracking.
codex AGENTS.md
AGENTS.md instructions for openai/codex, covering rust/codex-rs, the codex-core crate, code review rules, crate api surface and model visible context.
langchain AGENTS.md
AGENTS.md instructions for langchain-ai/langchain, covering global development guidelines for the langchain monorepo, corridor security analysis, project architecture and context, monorepo structure and development tools & commands.
vscode oss-third-party-notices.instructions.md
Instructions for microsoft/vscode, covering vs code oss third-party-notices pipeline, architecture, pipeline flow in ci, applying the notice (cutover) and fallback chain (never fail the build).
next.js AGENTS.md
Instructions for vercel/next.js, covering next.js development guide, codebase structure, monorepo overview, core package: packages/next and other important packages.