A video voice-over tool that creates spoken narration matched to the video's content and replaces the original audio track. Ark TTS is a text-to-speech service that turns written words into speech.
A video transcription and subtitle tool that extracts spoken words and precise timing, then creates JSON and SRT subtitle files. ASR means automatic speech recognition, or converting speech into text.
A speech-to-text service based on Doubao's LAS-ASR-PRO model. It can transcribe audio or video into timestamped sentence data, with optional speaker and speech analysis.
A skill for extracting audio from videos or audio files and splitting it into equal-length parts. It can convert formats such as WAV, MP3, and FLAC, and can connect input and output to TOS storage.
A Chinese-language video-editing skill for finding and extracting moments from long videos using natural-language requests, including scenes, highlights, people, or objects identified from a reference image.
A skill for creating short videos with ByteDance's Seedance model from text, images, or other reference media. Seedance is an AI video-generation model.
A skill for generating images with ByteDance's Seedream models from text or reference images. Seedream is an AI image-generation service supporting different visual styles and image shapes.
A text-to-speech tool that turns written text into playable audio using Doubao’s voice service. It can handle different voices and adjust speaking speed, pitch, and volume.
A speech-to-text tool using Volcengine BigModel ASR, a service that converts audio into written text. It supports Feishu voice messages, local audio, and audio URLs, with synchronous and asynchronous modes.
A video editing workflow that cuts and joins clips with FFmpeg, a command-line program for processing video, using timestamp data in a structured JSON file.
A tool for creating 5–8 second AI-generated opening videos with Seedance 2.0. These videos are intended to attract attention at the start of paid advertising material.
A Remotion workflow for adding karaoke-style lyrics to video or audio. Remotion is a tool for creating videos with code, while LRC is a timed-lyrics file format.
A workflow for making animated-story videos from a single sentence, including story expansion, visual assets, storyboards, short clips, and final assembly. It can also add voice-over and subtitles.
A skill for making music videos from an existing song and its lyrics, using speech recognition and generated images or video. It can align beginning and ending frames and optionally add permanent subtitles to the video.
An AI-based scoring tool for evaluating videos used in paid advertising. It looks at the opening three seconds, pacing, and the pattern of emotions in the video.
A Chinese-language skill for creating finished songs with ByteDance's GenSong service, including lyrics, audio, and structured song labels. It does not create music videos.
A workflow for correcting a few mispronounced words in an already-rendered speaking video by replacing only the affected audio segment. TTS, or text-to-speech, creates spoken audio from text.
Uma ferramenta que grava legendas SRT diretamente nos pixels de um vídeo. SRT é um formato de arquivo que guarda textos de legenda e seus tempos.
★not rated 5 5mo agoA105 tokens
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: