calesthio/generative-media-skills

Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants.

This repository also configures its own agents. See what generative-media-skills tells them →

171Stars on the repository
158Mods indexed here, across every type
2mo agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

hume-octave

121

calesthio/generative-media-skills

Skill Claude CodeCodex

Use Hume Octave for emotionally expressive speech and voice production: text-to-speech, voice design, voice cloning, voice conversion, streaming/realtime TTS, multilingual narration, dialogue continuity, timestamps/lip-sync, safety/rights review, and production QA.

not rated 171 +20 2mo ago A SkillSpector: warn 59 tokens original MIT

kokoro-tts

122

calesthio/generative-media-skills

Skill Claude CodeCodex

Local and self-hosted text-to-speech with Kokoro, the 82M-parameter open-weight (Apache 2.0) model by hexgrad. Use when synthesizing speech offline, on-device, in the browser, or on your own server without per-character API cost — for high-volume narration, audiobooks, privacy-constrained pipelines, and prototyping.…

not rated 171 +20 2mo ago A SkillSpector: pass 200 tokens original MIT

minimax-speech

123

calesthio/generative-media-skills

Skill Claude CodeCodex

Use this skill when producing speech, narration, dubbing, localization, advertising voice, voice-clone previews, or interactive voice with MiniMax speech/audio APIs. It covers MiniMax T2A HTTP, WebSocket streaming, async long-form TTS, Speech 2.8/2.6/02 model selection, system and custom voices, rapid voice cloning…

not rated 171 +20 2mo ago A SkillSpector: warn 105 tokens original MIT

qwen3-tts

124

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce text-to-speech with Alibaba/Qwen Qwen3-TTS through DashScope/Model Studio or open-weight Qwen3-TTS checkpoints. Use when an agent must choose Qwen3-TTS models, voices, realtime versus non-realtime synthesis, voice cloning, voice design, instruction/prosody controls, multilingual or dialect speech, audio…

not rated 171 +20 2mo ago A SkillSpector: warn 110 tokens original MIT

resemble-chatterbox

125

calesthio/generative-media-skills

Skill Claude CodeCodex

Use Resemble AI Chatterbox for text-to-speech and voice-cloning workflows, including local open-weight Chatterbox, Chatterbox Multilingual, Chatterbox Turbo, and Resemble-hosted Chatterbox API routes. Apply when an agent must choose a Chatterbox variant, prepare reference-voice inputs, control emotion/paralinguistic…

not rated 171 +20 2mo ago A SkillSpector: warn 110 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Production guidance for mainland-China Volcengine Doubao Speech text-to-speech. Use for TTS 1.0/2.0 selection, V3 bidirectional or unidirectional streaming, asynchronous long-text submit/query jobs, current voices and controls, timestamps and SSML caveats, expiring results, quota/error handling, authorized cloned…

not rated 171 +20 2mo ago A SkillSpector: pass 97 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Use when enhancing, upscaling, restoring, denoising, sharpening, deinterlacing, stabilizing, motion-deblurring, colorizing, SDR-to-HDR converting, or frame-rate/slow-motion converting video (and secondarily images) with Topaz Labs — whether cleaning up AI-generated video, compressed UGC, or old archival footage, or…

not rated 171 +20 2mo ago A SkillSpector: warn 174 tokens original MIT

alibaba-wan-video

128

calesthio/generative-media-skills

Skill Claude CodeCodex

Plan, implement, and review Alibaba Wan video generation using Alibaba Cloud Model Studio/DashScope hosted APIs or official Wan 2.1/2.2 open-weight checkpoints. Use for Wan text-to-video, image-to-video, first/last-frame or continuation, reference-to-video, video editing, speech-driven video, character animation…

not rated 171 +20 2mo ago A SkillSpector: pass 84 tokens original MIT

amazon-nova-reel

129

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce and operate Amazon Nova Reel video-generation jobs through Amazon Bedrock. Use for Nova Reel text-to-video, image-conditioned animation, automated or manual multi-shot storyboards, async S3 delivery, cost and approval gates, production continuity, output custody, and AWS-specific safety, privacy, IAM, and…

not rated 171 +20 2mo ago A SkillSpector: pass 69 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Generate, reference, and conversationally edit short videos with Google's Gemini Omni Flash through the Gemini Developer API or Gemini Enterprise Agent Platform. Use when a task specifically needs Gemini Omni video, multimodal reference roles, time-directed prompting, audio-aware generation, Interactions API state, or…

not rated 171 +20 2mo ago A SkillSpector: warn 81 tokens original MIT

google-veo

131

calesthio/generative-media-skills

Skill Claude CodeCodex

Direct production with Google DeepMind's Veo video-generation family across the Gemini Developer API and Google Cloud Gemini Enterprise Agent Platform (formerly Vertex AI). Use for Veo route and model selection, text/image/reference/first-last-frame/extension workflows, native-audio prompting, camera direction, async…

not rated 171 +20 2mo ago C 75 tokens original MIT

higgsfield-video

132

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce video (and its supporting stills) through Higgsfield (higgsfield.ai), a multi-model creative platform that fronts third-party video models (Kling, Veo, Seedance, Wan, MiniMax, and, until retirement, Sora) plus Higgsfield's own control layer — Soul / Soul ID image generation, the DoP camera-preset engine, and…

not rated 171 +20 2mo ago A SkillSpector: warn 183 tokens original MIT

kling-video

133

calesthio/generative-media-skills

Skill Claude CodeCodex

Plan, prompt, generate, reference, edit, motion-control, and quality-review videos with Kuaishou's direct Kling AI video platform and API, especially Kling VIDEO 3.0, 3.0 Omni, 3.0 Turbo, native audio, multi-shot, elements, start/end frames, and image/video references. Use for first-party Kling video production or API…

not rated 171 +20 2mo ago A SkillSpector: pass 105 tokens original MIT

ltx-2-video

134

calesthio/generative-media-skills

Skill Claude CodeCodex needs its repo

Generate and edit synchronized audio-video with LTX-2.3 using the official hosted LTX API or official local/open-weight LTX-2 repository. Use for text-to-video, image-to-video, audio-to-video, retake, extend, HDR conversion, local inference, checkpoint selection, quantization, and LTX-specific production planning.

not rated 171 +20 2mo ago A SkillSpector: pass 75 tokens original MIT

luma-ray-video

135

calesthio/generative-media-skills

Skill Claude CodeCodex

Direct and operate Luma Ray video generation through the current Luma Agents API and distinguish it from the consumer Luma App and legacy Dream Machine API. Use for Ray 3.2 text/image/multi-keyframe video, edits, extends, reframes, HDR/EXR, pricing and approvals, async delivery, migration, rights, privacy, and…

not rated 171 +20 2mo ago A SkillSpector: warn 79 tokens original MIT

midjourney-video

136

calesthio/generative-media-skills

Skill Claude CodeCodex

Plan and direct Midjourney Video V1 image-to-video work through the official website or Discord, with human operator handoff, motion prompts, start/end frames, loops, extensions, resolution and batch budgeting, privacy, rights, safety, and delivery QA. Use when creating Midjourney videos or when determining whether a…

not rated 171 +20 2mo ago A SkillSpector: pass 82 tokens original MIT

minimax-hailuo-video

137

calesthio/generative-media-skills

Skill Claude CodeCodex

Use MiniMax's first-party Hailuo video API safely and reproducibly across global and mainland-China platforms. Covers current text-to-video, image-to-video, first/last-frame, and subject-reference models; prompt and camera craft; region/account routing; exact cost approval; asynchronous jobs and callbacks; artifact…

not rated 171 +20 2mo ago C 128 tokens original MIT

moonvalley-marey

138

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce video with Moonvalley's Marey model family (Marey Realism v1.5) — a filmmaker-oriented, 1080p/24fps generative video model marketed as trained exclusively on licensed data. Use this skill when a task asks for Marey or Moonvalley specifically, when a brief demands brand-safe / legally reviewed AI video for…

not rated 171 +20 2mo ago A SkillSpector: pass 201 tokens original MIT

nvidia-cosmos-video

139

calesthio/generative-media-skills

Skill Claude CodeCodex

Select, run, and govern NVIDIA Cosmos world-video generation across Cosmos 3 Generator, Predict2.5, Transfer2.5, downloadable checkpoints, self-hosted NIMs, and hosted preview surfaces. Use for text/image/video-to-world, controlled world transfer, multiview or action-conditioned physical-AI video, local checkpoint…

not rated 171 +20 2mo ago A SkillSpector: warn 97 tokens original MIT

pika-video

140

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce short-form video with Pika (pika.art) — its effect-driven tools (Pikaffects, Pikadditions, Pikaswaps, Pikaframes, Pikascenes, Pikatwists) and audio-driven lip sync (Pikaformance), plus text-to-video and image-to-video. Use when a task calls for playful, stylized, meme-able social clips, or for applying…

not rated 171 +20 2mo ago A SkillSpector: warn 161 tokens original MIT

runway-video

141

calesthio/generative-media-skills

Skill Claude CodeCodex

Build and operate production-safe Runway API video generation, video editing, and character-performance workflows with native Runway models, explicit paid-call approval, duplicate-create protection, asynchronous task handling, secure media transfer, and evidence-aware governance.

not rated 171 +20 2mo ago A SkillSpector: warn 50 tokens original MIT

seedance-2-0

142

calesthio/generative-media-skills

Skill Claude CodeCodex

Direct ByteDance Dreamina Seedance 2.0 Standard, Fast, and Mini video production across BytePlus ModelArk and verified gateways. Use for text/image/reference-to-video, native synchronized audio and dialogue, multimodal image-video-audio reference, first/last-frame animation, video editing or extension, model/gateway…

not rated 171 +20 2mo ago A SkillSpector: pass 90 tokens original MIT

tencent-hunyuanvideo

143

calesthio/generative-media-skills

Skill Claude CodeCodex needs its repo

Generate and operate Tencent Hunyuan video through the managed TokenHub HY-Video-1.5 API or official local HunyuanVideo repositories. Use for text-to-video, image-to-video, hosted job lifecycle, pricing and region planning, local checkpoint selection, hardware and acceleration, prompt design, licensing, safety…

not rated 171 +20 2mo ago A SkillSpector: warn 77 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Select, integrate, and operate multi-model video-generation gateways — hosted inference aggregators such as fal.ai, Replicate, and WaveSpeed that expose many third-party video (and video-adjacent audio) models behind one account, one API surface, and one bill. Use when deciding whether to route video generation…

not rated 171 +20 2mo ago A SkillSpector: warn 168 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: