JayceVane

2 mods across 1 repository, 1 stars between them.

JayceVane/visual-understanding

Skill Claude CodeCodex

Multi-provider visual understanding tool for images, videos, and documents. Supports captioning, OCR, visual Q&A, document analysis, and object grounding (bounding-box localisation) through configurable providers (Zhipu GLM-V, OpenAI GPT-4o, Anthropic Claude, or any OpenAI-compatible endpoint). Use when the user wants…

1 15d ago A 103 tokens original MIT

JayceVane/visual-understanding

MCP server Claude CodeCodexCursor

Multi-provider visual understanding MCP server & CLI — caption, OCR, Q&A, grounding. Runs locally from the visual-understanding Python package.

1 15d ago A tokens not measured original MIT