JayceVane/visual-understanding
Skill Claude CodeCodex
Multi-provider visual understanding tool for images, videos, and documents. Supports captioning, OCR, visual Q&A, document analysis, and object grounding (bounding-box localisation) through configurable providers (Zhipu GLM-V, OpenAI GPT-4o, Anthropic Claude, or any OpenAI-compatible endpoint). Use when the user wants…
1 15d ago A 103 tokens
original MIT