kitlau86/agent-vision-mcp

An MCP server that gives non-vision LLMs the ability to "see" images. Plug in any OpenAI-compatible vision API (Gemini, Qwen-VL, OpenAI, or self-hosted) and your text-only model — like DeepSeek in Claude Code — can analyze screenshots, OCR text, read charts, and more.

These files are kitlau86/agent-vision-mcp's own configuration. They tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

10Stars on the repository
1Files it configures its agents with
436Tokens loaded in every session
1Agent configured

Instructions