calesthio/generative-media-skills

Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants.

This repository also configures its own agents. See what generative-media-skills tells them →

170Stars on the repository
158Mods indexed here, across every type
2mo agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce presenter-led AI avatar videos with Synthesia Studio and Synthesia API. Use for Synthesia-specific avatar video planning, avatar/voice/language selection, scripted scenes, template personalization, API job lifecycle, localization/dubbing, brand/training/product workflows, consent/rights/safety checks, artifact…

not rated 170 +20 2mo ago A SkillSpector: warn 73 tokens original MIT

tavus-replica-video

74

calesthio/generative-media-skills

Skill Claude CodeCodex

Use for producing Tavus AI-human videos and real-time avatar conversations with Tavus Faces/Replicas, PALs/Personas, async Video Generation, CVI conversations, consent-safe likeness workflows, webhooks, backgrounds, localization, and avatar QA.

not rated 170 +20 2mo ago A SkillSpector: warn 59 tokens original MIT

adobe-firefly-image

75

calesthio/generative-media-skills

Skill Claude CodeCodex

Use Adobe Firefly Services to generate, edit, expand, fill, match, composite, and upscale still images through the current REST APIs. Apply when selecting Firefly Image 5 versus Image 3/4 or custom models, implementing authenticated asynchronous image workflows, using style or structure references, building product…

not rated 170 +20 2mo ago C 97 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Plan, prompt, call, edit, iterate, and productionize Alibaba Cloud Model Studio image generation with current Wan 2.7 Image, Qwen-Image 2.0, and Z-Image Turbo models. Use for Alibaba/DashScope text-to-image, multi-reference generation, instruction editing, character-consistent image sets, typography and layout work…

not rated 170 +20 2mo ago A SkillSpector: pass 105 tokens original MIT

amazon-nova-canvas

77

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce and review production image-generation and image-editing workflows with Amazon Nova Canvas on Amazon Bedrock, including native InvokeModel payloads, safe authentication, validation, retries, cost controls, provenance, and lifecycle migration checks. Use when a task names Nova Canvas, amazon.nova-canvas-v1:0…

not rated 170 +20 2mo ago A SkillSpector: pass 100 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Plan, prompt, execute, troubleshoot, and quality-control Black Forest Labs FLUX image generation and editing across the BFL direct API and licensed local/open-weight deployments. Use for FLUX.2 model selection, text-to-image, single- or multi-reference editing, typography, exact-color work, mask-based…

not rated 170 +20 2mo ago B 111 tokens original MIT

bria-fibo-image

79

calesthio/generative-media-skills

Skill Claude CodeCodex

Build and operate rights-aware Bria FIBO and FIBO Lite image-generation workflows with structured prompts, reference images, asynchronous status handling, webhooks, cost gates, and safe artifact downloads. Use when a user asks for Bria/FIBO generation, refinement, inspiration, reproducibility, hosted API integration…

not rated 170 +20 2mo ago A SkillSpector: pass 99 tokens original MIT

bytedance-seedream

80

calesthio/generative-media-skills

Skill Claude CodeCodex

Build and operate production image generation and natural-language image editing with ByteDance Seedream through first-party Volcengine Ark (China) or BytePlus ModelArk (global), including model and region selection, multi-reference and grouped outputs, streaming, secure artifact handling, retries, cost controls…

not rated 170 +20 2mo ago A SkillSpector: warn 72 tokens original MIT

google-gemini-image

81

calesthio/generative-media-skills

Skill Claude CodeCodex

Plan, generate, edit, and quality-check still images with Google's Gemini native image models (Nano Banana), including Gemini 3.1 Flash Image, Gemini 3.1 Flash-Lite Image, Gemini 3 Pro Image, and legacy Gemini 2.5 Flash Image. Use for model and API-route selection, conversational editing, multi-reference composition…

not rated 170 +20 2mo ago A SkillSpector: warn 132 tokens original MIT

ideogram-image

82

calesthio/generative-media-skills

Skill Claude CodeCodex

Generate, remix, edit, inpaint, reframe, background-process, describe, layerize, and upscale images with the Ideogram Developer API. Use when Codex must integrate Ideogram 4.0 or 3.0, render reliable typography or structured layouts, use style or character references, build synchronous or webhook workflows, migrate…

not rated 170 +20 2mo ago A SkillSpector: warn 93 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Select, integrate, and operate multi-model image-generation gateways with model-specific schema discovery, version policy, asynchronous jobs, webhooks, spend approval, safe inputs and artifacts, data-governance review, and billing observability. Use for comparing or building against fal.ai, Replicate, or Together AI…

not rated 170 +20 2mo ago A SkillSpector: warn 100 tokens original MIT

kling-kolors-image

84

calesthio/generative-media-skills

Skill Claude CodeCodex

Plan, generate, edit, reference, and quality-control still images with Kling AI's hosted IMAGE surfaces or Kuaishou's open Kolors checkpoints. Use for Kling IMAGE 3.0/3.0 Omni/O1/2.1 web, official Kling CLI/MCP or Open Platform integration, and local Kolors text-to-image, image-to-image, IP-Adapter, ControlNet…

not rated 170 +20 2mo ago A SkillSpector: warn 133 tokens original MIT

leonardo-image

85

calesthio/generative-media-skills

Skill Claude CodeCodex

Create, edit, guide, upscale, and quality-control still images with Leonardo.Ai's official Production API, including native Lucid and Phoenix models, uploaded or generated references, image-to-image, ControlNet guidance, realtime-canvas inpainting, Pro and Universal upscalers, and custom Elements or models. Use for…

not rated 170 +20 2mo ago A SkillSpector: warn 118 tokens original MIT

luma-photon

86

calesthio/generative-media-skills

Skill Claude CodeCodex

Use Luma AI's first-party Photon and Photon Flash image API for text-to-image, image/style/character references, image modification, asynchronous generation, secure artifact handling, production prompting, cost control, and policy-aware delivery. Trigger for Luma Photon API integration or production work; do not use…

not rated 170 +20 2mo ago A SkillSpector: warn 84 tokens original MIT

midjourney-image

87

calesthio/generative-media-skills

Skill Claude CodeCodex

Plan, prompt, iterate, edit, migrate, and quality-check still-image work made with Midjourney's documented website and Discord interfaces. Use for Midjourney model and parameter selection, image/style/Omni references, Draft or Conversational workflows, typography-aware production, privacy and rights review, or…

not rated 170 +20 2mo ago A SkillSpector: warn 89 tokens original MIT

openai-gpt-image

88

calesthio/generative-media-skills

Skill Claude CodeCodex

Generate, edit, composite, stream, and production-review still images with OpenAI GPT Image models. Use when an agent must choose between GPT Image 2 and legacy GPT Image/DALL-E integrations, select the Image API or Responses API, build prompts and reference-image workflows, use masks or multiple inputs, preserve…

not rated 170 +20 2mo ago A SkillSpector: warn 124 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Design and produce raster images, native SVG artwork, brand-controlled visuals, and documented image edits with Recraft's API. Use for Recraft model and route selection, exact request construction, image-to-image, inpainting, outpainting, background work, vectorization, remix/exploration, palette and typography…

not rated 170 +20 2mo ago A SkillSpector: warn 80 tokens original MIT

runway-image

90

calesthio/generative-media-skills

Skill Claude CodeCodex

Generate, edit, and iterate still images with Runway's official API, especially native Gen-4 Image and Gen-4 Image Turbo reference workflows. Use when implementing Runway text-to-image, reference-driven image generation or natural-language image edits, task polling, secure artifact handling, production retries, cost…

not rated 170 +20 2mo ago A SkillSpector: warn 87 tokens original MIT

stability-ai-image

91

calesthio/generative-media-skills

Skill Claude CodeCodex

Operate Stability AI image generation, image-to-image, edit, control, background, and upscale APIs safely and reproducibly. Use when selecting or calling Stable Image Ultra, Core, Stable Diffusion 3.5, v2beta image edit/control services, or Stability open-weight image models; when debugging schemas, moderation, rate…

not rated 170 +20 2mo ago A SkillSpector: warn 95 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Generate and edit production images with xAI's first-party Grok Imagine API, including model selection, multiple references, synchronous and Batch API workflows, durable output handling, prompting, iteration, QA, cost and rate controls, privacy, safety, and rights review. Use for direct xAI image API integrations, not…

not rated 170 +20 2mo ago A SkillSpector: warn 78 tokens original MIT

amazon-rekognition

93

calesthio/generative-media-skills

Skill Claude CodeCodex

Use this skill when an agent needs production image or video understanding with Amazon Rekognition: labels, objects, scenes, OCR, moderation, image properties, Custom Labels or moderation adapters, stored-video analysis, conditional streaming-video workflows for existing eligible accounts, searchable media libraries…

not rated 170 +20 2mo ago A SkillSpector: warn 82 tokens original MIT

google-cloud-vision

94

calesthio/generative-media-skills

Skill Claude CodeCodex

Use this skill when an agent needs Google Cloud Vision API for still-image understanding: labels, object localization, OCR, document text, SafeSearch, image properties, crop hints, web detection, batch annotation, Cloud Storage based pipelines, confidence evaluation, privacy, quotas, cost, and QA. Do not use it for…

not rated 170 +20 2mo ago A SkillSpector: pass 96 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Production guidance for Kling AI Open Platform Advanced Lip-Sync. Use for identifying/selecting one face in an existing video, assigning URL/Base64 or Kling TTS audio, cropping and inserting audio in milliseconds, mixing original sound, creating and monitoring tasks, downloading expiring results, consent/privacy…

not rated 170 +20 2mo ago A SkillSpector: pass 87 tokens original MIT

sync-labs-lipsync

96

calesthio/generative-media-skills

Skill Claude CodeCodex

Generate AI lip sync and visual dubbing with Sync Labs (sync.so) — choose the right model (sync-3, lipsync-2, lipsync-2-pro, lipsync-1.9.0-beta, react-1), build the POST /v2/generate request, host inputs, handle async jobs via polling or webhooks, run batch dubbing/localization, and review output for sync accuracy and…

not rated 170 +20 2mo ago A SkillSpector: warn 156 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: