Plan and integrate ShengShu Vidu Open Platform video generation with current model/mode selection, reference consistency, pricing, exact approval, async task handling, and safe artifact custody. Use for Vidu API text-to-video, image-to-video, start/end-frame, multi-subject reference, media-reference, or Q2 multi-frame…
Produce, edit, extend, and govern short videos with xAI's direct Grok Imagine Video API. Use for text-to-video, image-to-video, multi-reference video, natural-language video edits, video continuation, exact media-cost approval, asynchronous request recovery, Files/Batch integration, moderation review, and API privacy…
Use when an agent must make TwelveLabs do real video-understanding work: indexing footage, semantic/visual search across an archive, generating descriptions, summaries, chapters, highlights, and tags from video, producing multimodal embeddings, or wiring TwelveLabs into a media-production pipeline (NLE panels…
Build, configure, and ship production voice agents on the ElevenLabs Agents platform (branded "ElevenAgents," formerly "Conversational AI"). Use when an agent must design a spoken-conversation system prompt, pick an LLM and TTS voice/model for a real-time voice bot, wire client/server/system tools and knowledge-base…
Build and evaluate low-latency spoken, multimodal, and translation experiences with Google Gemini Live API and Gemini Enterprise Agent Platform Live API. Use when a media-production or voice-agent workflow needs real-time audio input/output, barge-in, voice configuration, live transcription, live translation…
Build production OpenAI Realtime voice agents and low-latency spoken interactions with live audio sessions, WebRTC or WebSocket transport, voice activity detection, tool/function calling, prompt design, logging, consent, privacy, safety, latency, and cost controls. Use for speech-to-speech agents and live voice UX; do…
Plan and scope production work with Odyssey's real-time interactive video world models (odyssey.ml — Odyssey-1/Odyssey-2/Starchild-1/Agora-1). Use when a request involves generating video that a user steers live with keyboard, text, speech, or controller input (streamed frames that respond in the moment), or when…
Generate persistent, explorable 3D worlds/environments with Marble by World Labs — from text, a single image, multiple images, video, or a 360 panorama, plus coarse-structure blocking with Chisel. Use this skill when a task needs a navigable 3D scene, VR/immersive backdrop, virtual-production environment, previz set…
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: