multimodal skills

152 tagged multimodal, measured the same way as everything else here.

Browse within: vision 21ocr 19GLM 18agentic-workflow 17ai-glasses 17clawdbot 17edge-ai 17meta-evolution 17Distributed Training 14GRPO 14Multi-Agent 14Post-Training 14agentic-rl 14self-hosted 14

llamaindex

01

davila7/claude-code-templates

Skill Claude CodeCodex

Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal support. Use for document Q&A, chatbots, knowledge retrieval, or building RAG pipelines. Best for data-centric LLM…

30k +27 today A 70 tokens original MIT

pixelbrowse

02

StarTrail-org/PixelRAG

Skill Claude CodeCodex

Screenshot and visually read any web page or document using pixelshot. Use instead of fetching raw HTML when you need to see what a page looks like, read visual content (charts, diagrams, infographics), check layouts, or verify UI. Triggers: "look at this page", "screenshot", "what does this site look like", "check…

9.8k 2d ago A 90 tokens original Apache-2.0

typescript

03

genkit-ai/genkit

Skill Claude CodeCodex

TypeScript coding conventions, best practices, and patterns for writing clean, maintainable code.

6.4k +2 today A 20 tokens original Apache-2.0

python-expert

04

genkit-ai/genkit

Skill Claude CodeCodex

Conventions for clean, idiomatic Python. Load whenever you read, edit, or write Python source files.

6.4k +2 today A 26 tokens original Apache-2.0

genkit-ai/genkit

Skill Claude CodeCodex

Best practices for authoring Genkit tooling, including CLI commands and MCP server tools. Covers naming conventions, architectural patterns, and consistency guidelines.

6.4k +2 today A 35 tokens original Apache-2.0

create-pr

06

NVIDIA-NeMo/DataDesigner

Skill Claude CodeCodex

Create a GitHub PR with a well-formatted description matching the repository PR template (flat Changes by default; optional Added/Changed/Removed/Fixed grouping).

2.2k 2d ago A 35 tokens original Apache-2.0

datadesigner-docs

07

NVIDIA-NeMo/DataDesigner

Skill Claude CodeCodex

Maintain the NeMo Data Designer Fern docs site under fern/. Use for any documentation change. Triggered by: "edit docs", "add doc page", "update docs", "rename page", "fix broken link", "add redirect", "preview docs", "publish docs", "regenerate notebooks", "update dev note", any request that touches fern/.

2.2k 2d ago A 80 tokens original Apache-2.0

review-code

08

NVIDIA-NeMo/DataDesigner

Skill Claude CodeCodex

Perform a thorough code review of the current branch or a GitHub PR by number.

2.2k 2d ago A 20 tokens original Apache-2.0

vision-skills

09

Anionex/agent-vision-toolkit

Skill Claude CodeCodex

Local vision CLIs: glance (describe/ask/OCR an image), ground (locate a target, pixel box), detect (element inventory), trace (image to SVG geometry), crop (cut a pixel box to a file), and scripts/htmlshot.py (HTML file to image). Use for any task involving an image — questions, text, splitting and transcribing long…

1.1k 5d ago A 132 tokens original MIT

debug-hang

10

redai-infra/Relax

Skill Claude CodeCodex

A troubleshooting workflow for Ray, a system that runs machine-learning jobs across multiple computers or GPUs, when a distributed training job stops making progress.

580 3d ago A 67 tokens original Apache-2.0

redai-infra/Relax

Skill Claude CodeCodex

Integrate a new NVIDIA NeMo Gym environment into Relax as a three-step recipe. Use when adding or debugging a recipe under examples/nemogymagentic/recipes; covers data preparation, a local private Gym service, direct Ray training launch, verifier validation, callback networking, lifecycle cleanup, and failure triage.

580 3d ago A 73 tokens original Apache-2.0

verl-to-relax

12

redai-infra/Relax

Skill Claude CodeCodex

Migrate RL training recipes from verl to Relax framework. Use when user wants to port reward functions, tool environments, training scripts, or any recipe code from the verl (volcengine/verl) codebase to Relax. Handles reward, rollout, tool/env, dataset, and launch script conversion. Supports both colocate (default)…

580 3d ago A 78 tokens original Apache-2.0

docs-governance

13

clawdotnet/openclaw.net

Skill Claude CodeCodex

Organizes repository documentation and keeps new docs in the correct location.

487 4d ago A 18 tokens original MIT

clawdotnet/openclaw.net

Skill Claude CodeCodex

Extract structured insight briefs from community research transcripts or notes. Produces pain points, stakeholder needs, opportunity maps, risks, and follow-up questions. Requires human review before publication.

487 4d ago A 41 tokens original MIT

history-explorer

15

clawdotnet/openclaw.net

Skill Claude CodeCodex

Inspect recent session conversation turns and meta-run history. Returns structured JSON summaries of turn counts, role distribution, tool usage, meta-skill executions, and co-occurrence patterns. Use this before running meta-skill-creator or when the user asks about recent history, past tool usage, or session activity.

487 4d ago A 65 tokens original MIT

glm-image-gen

16

zai-org/GLM-skills

Skill Claude CodeCodex

Official skill for generating high-quality images from text prompts using ZhiPu GLM-Image API. Excellent at scientific illustrations, high-quality portraits, social media graphics, and commercial posters. Supports multiple aspect ratios, HD quality, and watermark control. Use this skill when the user wants to generate…

468 4mo ago A 80 tokens original Apache-2.0

glmv-stock-analyst

17

zai-org/GLM-skills

Skill Claude CodeCodex

A stock-analysis workflow for Hong Kong, mainland Chinese, and United States shares. It combines company information, price charts, trading activity, news, and wider economic factors into a report.

468 4mo ago A 178 tokens original Apache-2.0

zai-org/GLM-skills

Skill Claude CodeCodex

Frontend visual replication skill. Explores a target website’s publicly visible pages via Playwright MCP or agent-browser, captures screenshots and layout information, then generates a static or client-side frontend replica that approximates the original’s visual appearance and page structure. This skill replicates…

468 4mo ago A 151 tokens original Apache-2.0

vision

20

xiincs/claude-code-vision-skill

Skill Claude CodeCodex

Call vision models (Doubao, Qwen, DeepSeek, OpenAI) to analyze images. Use when you need to understand screenshots, UI layouts, diagrams, or any image content. Supports png/jpg/webp/gif.

171 7d ago A 49 tokens original MIT

cerul

21

cerul-ai/cerul

Skill Claude CodeCodex

Use Cerul when a user needs cited evidence from video or long-form media, or asks what was said, shown, or presented.

156 2d ago A 30 tokens original Apache-2.0

aesthetic_drawing

22

lcqysl/GEMS

Skill Claude CodeCodex

This skill is designed to rewrite user prompts to align with the expert-level aesthetic standards. It transforms simple descriptions into multi-dimensional, professional-grade artistic instructions that maximize scores across all fine-grained aesthetic attributes.

145 5mo ago A 0 tokens

creative_drawing

23

lcqysl/GEMS

Skill Claude CodeCodex

This skill should be triggered when the user's request implies a need for imagination, artistic flair, conceptual depth, or unconventional visual storytelling.

145 5mo ago A 0 tokens

spatial

24

lcqysl/GEMS

Skill Claude CodeCodex

This skill should be triggered when the user's request involves multiple objects, complex scene arrangements, or specific physical relationships between elements.

145 5mo ago A 0 tokens