vision
01Skill Claude CodeCodex
Locally-runnable multimodal vision & speech for text-only LLM agents (Codex, Claude Code, Cursor, Cline, Gemini CLI, etc.). Use when the user asks to read, analyze, transcribe, or summarize images, screenshots, PDFs, charts, UI captures, videos, subtitles, audio, or media URLs — or when image input fails with "does…
not rated 1 20d ago C 179 tokens
original MIT