calesthio/generative-media-skills

Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants.

This repository also configures its own agents. See what generative-media-skills tells them →

170Stars on the repository
158Mods indexed here, across every type
2mo agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

calesthio/generative-media-skills

Skill Claude CodeCodex

Use DeepMotion Animate 3D for markerless video-to-3D human motion capture through the cloud portal, sales-gated API, or real-time SDK; covers capture planning, body/hand/face/multi-person limits, custom characters, retargeting, refinement, exports, DCC/game-engine handoff, QA, pricing, privacy, consent, and rights…

not rated 170 +20 2mo ago A SkillSpector: pass 85 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Use this skill when planning, capturing, processing, validating, troubleshooting, or exporting markerless human motion capture with Move AI products, including Move One single-camera workflows and Genesis multi-camera studio systems. It covers provider selection, setup, calibration, wardrobe and lighting constraints…

not rated 170 +20 2mo ago A SkillSpector: pass 122 tokens original MIT

ace-step

99

calesthio/generative-media-skills

Skill Claude CodeCodex

Use ACE-Step and ACE-Step 1.5 for local or hosted AI music generation, including text-to-music, lyrics-to-song, instrumental beds, covers, repainting, stem/track extraction, track completion, LoRA personalization, REST/Python/Gradio workflows, rights review, and music integration for video, ads, games, and social…

not rated 170 +20 2mo ago A SkillSpector: pass 76 tokens original MIT

elevenlabs-music

100

calesthio/generative-media-skills

Skill Claude CodeCodex

Generate and iterate music with ElevenLabs Eleven Music for production deliverables. Use when planning, prompting, API-calling, editing, inpainting, reviewing, or licensing AI-generated songs, instrumental beds, ad music, soundtrack cues, video-to-music scores, stems, or music assets using the ElevenLabs Music API or…

not rated 170 +20 2mo ago A SkillSpector: pass 75 tokens original MIT

google-lyria

101

calesthio/generative-media-skills

Skill Claude CodeCodex

Use Google Lyria music generation models through Google Cloud / Vertex AI / Gemini Enterprise Agent Platform for production music beds, songs, vocal tracks, image-conditioned music, short-form social audio, ad music, and video scoring. Covers Lyria 2, Lyria 3 Clip, and Lyria 3 Pro model selection, prompt construction…

not rated 170 +20 2mo ago A SkillSpector: pass 90 tokens original MIT

minimax-music

102

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce music with MiniMax Music 2.6 and MiniMax cover/lyrics APIs for songs, instrumentals, AI-generated lyrics, reference-audio covers, video/social/ad soundtracks, artifact custody, rights checks, and production QA.

not rated 170 +20 2mo ago A SkillSpector: warn 53 tokens original MIT

suno-music

103

calesthio/generative-media-skills

Skill Claude CodeCodex

Generate music with Suno (v5 / v5.5 era, 2026) and advise a production team on what they may legally do with the output. Use this skill when a user wants to create, extend, remaster, or stem-separate a track in Suno; craft Suno prompts (style field, lyrics/metatags, exclude-styles, personas, custom voices); choose a…

not rated 170 +20 2mo ago A SkillSpector: pass 176 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce, direct, generate, QA, and integrate non-speech audio with ElevenLabs Text to Sound Effects / Sound Effects API. Use when an agent needs custom SFX, foley, ambience, loops, stingers, impacts, UI sounds, musical one-shots, trailer braams, or sound-design layers for video, ads, social edits, games, apps…

not rated 170 +20 2mo ago A SkillSpector: warn 94 tokens original MIT

stable-audio

105

calesthio/generative-media-skills

Skill Claude CodeCodex

Use for Stability AI Stable Audio production work: selecting Stable Audio hosted API or open-weight models, generating or editing music, loops, sound effects, foley, beds, stingers, and sonic-branding audio from text or source audio, planning rights-safe uploads, setting model parameters, polling asynchronous jobs…

not rated 170 +20 2mo ago A SkillSpector: warn 73 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Operate AudioShake's cloud source-separation service (developer.audioshake.ai) to split a recording into stems. Use when a task involves isolating vocals, drums, bass, guitar, piano, keys, strings, or winds from music; separating dialogue / music / effects (DME) for film-TV post-production, localization, and dubbing…

not rated 170 +20 2mo ago A SkillSpector: warn 211 tokens original MIT

azure-speech

107

calesthio/generative-media-skills

Skill Claude CodeCodex

Use Microsoft Azure Speech in Foundry Tools for media-production speech workflows: speech-to-text, fast and batch transcription, diarization, captions/subtitles, real-time transcription, text-to-speech neural and HD voices, SSML, custom/personal voices, text-to-speech avatars, speech/video translation, localization…

not rated 170 +20 2mo ago A SkillSpector: warn 81 tokens original MIT

deepgram-speech

108

calesthio/generative-media-skills

Skill Claude CodeCodex

Use for Deepgram speech and voice production workflows: speech-to-text transcription, live captions, diarization, audio intelligence, Aura text-to-speech, Flux and Voice Agent live audio, model selection, cost/limits/privacy checks, artifact custody, and production QA.

not rated 170 +20 2mo ago A SkillSpector: warn 58 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Use for ElevenLabs provider-specific dubbing, localization, voice changer, speech-to-speech, and voice isolation/enhancement workflows. Applies when an agent must prepare source media, choose ElevenLabs dubbing versus voice conversion, manage language/speaker/timing decisions, handle voice identity and consent, call…

not rated 170 +20 2mo ago A SkillSpector: pass 103 tokens original MIT

google-cloud-speech

110

calesthio/generative-media-skills

Skill Claude CodeCodex

Use this skill when a media-production agent needs Google Cloud speech and voice services for transcription, captions/subtitles, long-form audio analysis, live caption planning, Text-to-Speech narration, Chirp 3 / Gemini-TTS voice selection, consented custom voice workflows, dubbing/localization planning…

not rated 170 +20 2mo ago A SkillSpector: pass 84 tokens original MIT

nvidia-speech-nim

111

calesthio/generative-media-skills

Skill Claude CodeCodex

Use NVIDIA Speech NIM microservices for speech production and voice workflows, including self-hosted ASR/STT, TTS, text translation, speech-to-speech pipeline design, Riva Python client integration, deployment planning, GPU/runtime sizing, licensing/privacy review, and production QA for transcription, captioning…

not rated 170 +20 2mo ago A SkillSpector: warn 79 tokens original MIT

openai-audio

112

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce and understand audio with OpenAI request-based audio APIs and audio-capable chat models, including text-to-speech, transcription, translation, multimodal audio input/output, model routing, prompt and performance direction, artifact custody, approval gates, and safety/rights review. Use for non-realtime OpenAI…

not rated 170 +20 2mo ago A SkillSpector: warn 81 tokens original MIT

amazon-transcribe

113

calesthio/generative-media-skills

Skill Claude CodeCodex

Use Amazon Transcribe for AWS-based speech-to-text production: batch S3 transcription, real-time streaming, captions/subtitles, diarization, channel identification, custom vocabularies, vocabulary filters, language identification, PII/PHI handling, toxicity detection, Call Analytics, Medical, and secure S3/IAM/KMS…

not rated 170 +20 2mo ago B 71 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Use this skill when an agent needs AssemblyAI for speech-to-text or speech-understanding work in media production, including pre-recorded, synchronous short-file, and real-time streaming transcription; speaker diarization or speaker identification; captions and subtitles; timestamps; language detection…

not rated 170 +20 2mo ago A SkillSpector: warn 125 tokens original MIT

elevenlabs-scribe

115

calesthio/generative-media-skills

Skill Claude CodeCodex

Use ElevenLabs Scribe for speech-to-text production workflows: transcribing audio or video files, diarization and speaker/channel handling, word/character timestamps, captions/subtitles, keyterm prompting, entity detection/redaction, webhooks, realtime STT boundaries, pricing/limits, privacy/retention, consent…

not rated 170 +20 2mo ago A SkillSpector: warn 88 tokens original MIT

amazon-polly

116

calesthio/generative-media-skills

Skill Claude CodeCodex

Use Amazon Polly for production text-to-speech work: selecting Standard, Neural, Long-form, or Generative engines and compatible voices; authoring SSML; creating speech marks for captions, word highlighting, or lip-sync; managing pronunciation lexicons; running synchronous, streaming, or asynchronous S3-backed…

not rated 170 +20 2mo ago A SkillSpector: warn 92 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Production guidance for international BytePlus Seed Speech text-to-speech. Use for selecting TTS 1.0 versus 2.0, bidirectional or unidirectional streaming, current voices and languages, prompt/prosody controls, subtitle timing validation, billing and concurrency, authorized replicated voices, privacy, error handling…

not rated 170 +20 2mo ago A SkillSpector: pass 91 tokens original MIT

cartesia-sonic

118

calesthio/generative-media-skills

Skill Claude CodeCodex

Use Cartesia Sonic and related Cartesia voice APIs for production speech: text-to-speech, realtime WebSocket TTS, voice selection, instant and professional voice cloning, pronunciation/language/emotion controls, voice localization, voice changer, pricing/concurrency planning, privacy/security review, and QA for…

not rated 170 +20 2mo ago A SkillSpector: warn 80 tokens original MIT

elevenlabs-tts

119

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce, direct, integrate, and quality-control ElevenLabs text-to-speech for narration, character performance, multilingual media, long-form voiceover, and streamed or batch AI-content audio. Use when selecting ElevenLabs models or voices; preparing spoken text; controlling pronunciation, pacing, emotion, and…

not rated 170 +20 2mo ago A SkillSpector: warn 117 tokens original MIT

fish-audio-tts

120

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce speech and clone voices with Fish Audio — its hosted TTS API (S2.1-Pro / S2-Pro / S1 model lineup, REST + WebSocket streaming, instant and persistent voice cloning) and its open-weight OpenAudio S1-mini / Fish-Speech models for self-hosting. Use this skill when an agent must generate narration or dialogue…

not rated 170 +20 2mo ago A SkillSpector: warn 163 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: