calesthio/generative-media-skills

Research-backed agent skills and tools for premium image, video, audio, voice, and generative media production across AI coding assistants.

This repository also configures its own agents. See what generative-media-skills tells them →

171Stars on the repository
158Mods indexed here, across every type
2mo agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

vidu-video

145

calesthio/generative-media-skills

Skill Claude CodeCodex

Plan and integrate ShengShu Vidu Open Platform video generation with current model/mode selection, reference consistency, pricing, exact approval, async task handling, and safe artifact custody. Use for Vidu API text-to-video, image-to-video, start/end-frame, multi-subject reference, media-reference, or Q2 multi-frame…

not rated 171 +20 2mo ago A SkillSpector: warn 93 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Produce, edit, extend, and govern short videos with xAI's direct Grok Imagine Video API. Use for text-to-video, image-to-video, multi-reference video, natural-language video edits, video continuation, exact media-cost approval, asynchronous request recovery, Files/Batch integration, moderation review, and API privacy…

not rated 171 +20 2mo ago A SkillSpector: warn 120 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Use when an agent must make TwelveLabs do real video-understanding work: indexing footage, semantic/visual search across an archive, generating descriptions, summaries, chapters, highlights, and tags from video, producing multimodal embeddings, or wiring TwelveLabs into a media-production pipeline (NLE panels…

not rated 171 +20 2mo ago A SkillSpector: pass 167 tokens original MIT

elevenlabs-agents

148

calesthio/generative-media-skills

Skill Claude CodeCodex

Build, configure, and ship production voice agents on the ElevenLabs Agents platform (branded "ElevenAgents," formerly "Conversational AI"). Use when an agent must design a spoken-conversation system prompt, pick an LLM and TTS voice/model for a real-time voice bot, wire client/server/system tools and knowledge-base…

not rated 171 +20 2mo ago A SkillSpector: warn 153 tokens original MIT

gemini-live-audio

149

calesthio/generative-media-skills

Skill Claude CodeCodex

Build and evaluate low-latency spoken, multimodal, and translation experiences with Google Gemini Live API and Gemini Enterprise Agent Platform Live API. Use when a media-production or voice-agent workflow needs real-time audio input/output, barge-in, voice configuration, live transcription, live translation…

not rated 171 +20 2mo ago A SkillSpector: warn 95 tokens original MIT

hume-evi

150

calesthio/generative-media-skills

Skill Claude CodeCodex

Build production realtime voice agents with Hume's Empathic Voice Interface (EVI): speech-to-speech sessions, empathic/prosodic response design, EVI 3 versus EVI 4-mini selection, WebSocket and SDK integration, tools/function calling, interruption/barge-in, turn-taking, transcripts/audio artifacts, pricing/limits…

not rated 171 +20 2mo ago A SkillSpector: warn 85 tokens original MIT

openai-realtime-voice

151

calesthio/generative-media-skills

Skill Claude CodeCodex

Build production OpenAI Realtime voice agents and low-latency spoken interactions with live audio sessions, WebRTC or WebSocket transport, voice activity detection, tool/function calling, prompt design, logging, consent, privacy, safety, latency, and cost controls. Use for speech-to-speech agents and live voice UX; do…

not rated 171 +20 2mo ago A SkillSpector: warn 91 tokens original MIT

calesthio/generative-media-skills

Skill Claude CodeCodex

Plan and scope production work with Odyssey's real-time interactive video world models (odyssey.ml — Odyssey-1/Odyssey-2/Starchild-1/Agora-1). Use when a request involves generating video that a user steers live with keyboard, text, speech, or controller input (streamed frames that respond in the moment), or when…

not rated 171 +20 2mo ago A SkillSpector: warn 163 tokens original MIT

world-labs-marble

153

calesthio/generative-media-skills

Skill Claude CodeCodex

Generate persistent, explorable 3D worlds/environments with Marble by World Labs — from text, a single image, multiple images, video, or a 360 panorama, plus coarse-structure blocking with Chisel. Use this skill when a task needs a navigable 3D scene, VR/immersive backdrop, virtual-production environment, previz set…

not rated 171 +20 2mo ago A SkillSpector: pass 204 tokens original MIT

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: