vision skills

44 tagged vision, measured the same way as everything else here.

Browse within: Multimodal 21ocr 20GLM 17image 6

glm-image-gen

01

zai-org/GLM-skills

Skill Claude CodeCodex

Official skill for generating high-quality images from text prompts using ZhiPu GLM-Image API. Excellent at scientific illustrations, high-quality portraits, social media graphics, and commercial posters. Supports multiple aspect ratios, HD quality, and watermark control. Use this skill when the user wants to generate…

468 4mo ago A 80 tokens original Apache-2.0

glmv-stock-analyst

02

zai-org/GLM-skills

Skill Claude CodeCodex

A stock-analysis workflow for Hong Kong, mainland Chinese, and United States shares. It combines company information, price charts, trading activity, news, and wider economic factors into a report.

468 4mo ago A 178 tokens original Apache-2.0

zai-org/GLM-skills

Skill Claude CodeCodex

Frontend visual replication skill. Explores a target website’s publicly visible pages via Playwright MCP or agent-browser, captures screenshots and layout information, then generates a static or client-side frontend replica that approximates the original’s visual appearance and page structure. This skill replicates…

468 4mo ago A 151 tokens original Apache-2.0

luma-vision

04

JochenYang/luma-mcp

Skill Claude CodeCodex

A skill for analyzing images with several external vision models. It is activated only when a user starts a request with /skill luma-vision and includes an image.

113 23d ago A 37 tokens original MIT

delphi-uses-graph

05

MaxiDonkey/DelphiAnthropic

Skill Claude CodeCodex

Analyzes a Delphi / Object Pascal codebase to extract unit-level uses dependencies. Use when the user uploads a zip / archive of a Delphi project (or a folder of .pas / .dpr / .dpk files) and asks for a dependency graph, architecture map, cycle detection, fan-in / fan-out coupling analysis, or wants to know how units…

60 3mo ago A 124 tokens original MIT

asdhabdua/bilibili-video-notes-skill

Skill Claude CodeCodex

A tool that turns educational or lecture videos from Bilibili, a Chinese video-sharing site, into DOCX study notes. It uses subtitles, screenshots, text recognition from images, and visual review to build the notes.

44 1mo ago A 58 tokens

media-tools

07

MJorgin/dsh-media-skills

Skill Claude CodeCodex

A guide and script for generating images, illustrations, avatars, backgrounds, and banners through configured image-generation services. It uses available API keys from environment or local secret files.

19 8d ago A 77 tokens original MIT

vision-review

08

MJorgin/dsh-media-skills

Skill Claude CodeCodex

A tool for reading images and checking screenshots, including text, layout, visual defects, watermarks, and logos.

19 8d ago A 150 tokens original MIT

dsv-bridge

09

menghuanshiguang/dsv-bridge-skill

Skill Claude CodeCodex

An image-understanding bridge for local AI systems that cannot read or interpret pictures. It sends images to DeepSeek’s vision mode and returns a text answer.

11 28d ago A 90 tokens original MIT

deepseek-vision

10

reF0o0/deepseek-vision-skill

Skill Claude CodeCodex

MUST use when the user sends or asks about images, photos, screenshots, pictures, audio, video, or mixed media documents, including requests to OCR/read text from an image. Route all media through Xiaomi MiMo V2.5 (mimo-v2.5) and mimo-v2.5-asr via scripts/mimo.py; never use local OCR, viewimage, native vision…

5 15d ago A 108 tokens

image-understanding

11

famclaw/famclaw

Skill Claude CodeCodex

Enable image understanding capabilities for describing/analyzing images, reading text, and answering questions about visual content.

5 7d ago A 25 tokens AGPL-3.0

visionbuddy

12

FanZhangnan/VisionBuddy

Skill Claude CodeCodex

Analyze local screenshots and image files with VisionBuddy through the installed WorkBuddy CodeBuddy CLI using free Hy3 or explicitly authorized DeepSeek V4 Flash. Use when Codex needs OCR, layout inspection, UI/design analysis, image description, visual comparison, or evidence from PNG, JPEG, WebP, GIF, or BMP files…

2 25d ago A 87 tokens original MIT

image-reading

13

jing-hy/picturereader-zcode

Skill Claude CodeCodex

Read and understand images like a multimodal model using the picturereader tools (imagescan / imageocr / imagesample). Applies a verified 5-step workflow (global tone → find subjects → verify text → judge material → synthesize) guided by grounded principles and cross-image insights. Use whenever you need to look at an…

2 15d ago A 72 tokens original MIT archived