JayceVane/visual-understanding

Multi-provider visual understanding MCP server + CLI — caption, OCR, Q&A, grounding (Zhipu GLM-V / OpenAI / Anthropic / any OpenAI-compatible endpoint)

1Stars on the repository
2Mods indexed here, across every type
15d agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

JayceVane/visual-understanding

Skill Claude CodeCodex

Multi-provider visual understanding tool for images, videos, and documents. Supports captioning, OCR, visual Q&A, document analysis, and object grounding (bounding-box localisation) through configurable providers (Zhipu GLM-V, OpenAI GPT-4o, Anthropic Claude, or any OpenAI-compatible endpoint). Use when the user wants…

1 15d ago A 103 tokens original MIT