vision
01Skill Claude CodeCodex
Locally-runnable multimodal vision & speech for text-only LLM agents (Codex, Claude Code, Cursor, Cline, Gemini CLI, etc.). Use when the user asks to read, analyze, transcribe, or summarize images, screenshots, PDFs, charts, UI captures, videos, subtitles, audio, or media URLs — or when image input fails with "does…
1 17d ago C 179 tokens
original MIT