Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add calesthio/generative-media-skills --skill amazon-transcribegit clone --depth 1 https://github.com/calesthio/generative-media-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/calesthio/generative-media-skills/amazon-transcribe)<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/amazon-transcribe"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/amazon-transcribe/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/amazon-transcribe"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/amazon-transcribe.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00071 | $0.05275 |
| Opus 5 | $0.00036 | $0.02638 |
| Sonnet 5 | $0.00014 | $0.01055 |
| Haiku 4.5 | $0.00007 | $0.00528 |
Grade B, and why
amazon-transcribe scanned grade B with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Strips warnings and disclaimersmediumAnti-refusal
Omitting safety caveats hides risk from the user and is a common jailbreak preamble.
- Batch language identification plus content redaction has a critical boundary: if the audio contains languages other than supported redaction languages, only supported-language content is redacted and other languages ar How it starts
The opening of the file, as written. The whole thing — 265 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Amazon Transcribe
Use this skill when an agent needs to plan, run, review, or troubleshoot Amazon Transcribe as a speech-to-text provider inside an AWS-controlled workflow. Treat Amazon Transcribe as an AWS custody option first: it is strongest when the media already belongs in S3, the organization needs IAM/KMS/CloudTrail governance, captions must be generated from batch files, or real-time transcripts must flow through AWS streaming endpoints.
Do not use this skill as a generic "best transcription model" ranking. Compare accuracy, latency, cost, language coverage, and custody requirements against the actual job. For healthcare, call-center analytics, and sensitive audio, separate the general Transcribe path from the Medical and Call Analytics products before choosing.
Facts below were verified against AWS documentation on 2026-07-10 unless a line says otherwise.
Choose the right Amazon Transcribe surface
Documented fact: Amazon Transcribe is an automatic speech recognition service that converts audio to text. It can be used through asynchronous batch jobs for media files or real-time streaming for live audio. Sources: What is Amazon Transcribe, Streaming audio, StartTranscriptionJob API.
Use this decision map:
- Batch transcription: use for files in Amazon S3, offline transcripts, subtitles, captions, podcast/video backlogs, archives, QA review, and workflows that can wait for job completion. Batch requires the media file to be in S3 before
StartTranscriptionJob. - Streaming transcription: use for live captions, call assistants, meetings, voice input, or low-latency partial/final transcripts. Streaming uses bidirectional HTTP/2 or WebSockets and requires a language choice or language identification, media encoding, and sample rate. Source: StartStreamTranscription API.
- Call Analytics: use only when the media is a customer-agent or sales/support call and the desired output includes call-specific insights such as turn-based output, sentiment, interruptions, non-talk time, talk speed, categories, summaries, or action items. Do not use it as a generic captioning path. Sources: Call Analytics, post-call output.
- Medical: use only for medical-related speech such as clinician dictation, telemedicine, or clinician-patient conversations. Amazon Transcribe Medical is available for batch and streaming, is US English (
en-US) only, and AWS says it is not a substitute for professional medical advice, diagnosis, or treatment; patient-care uses require trained human review. Source: Amazon Transcribe Medical.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 265 lines · 71 tokens per session scan B 42306913b25e
amazon-transcribe is a skill published in the GitHub repository calesthio/generative-media-skills (170 stars, last pushed 2mo ago), licensed MIT. It adds 71 tokens to every session and 5,275 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it B with 1 finding (strips warnings and disclaimers). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
cliptalk-cover-director
Produces evidence-backed cover candidates and reviewable cover variants for a ClipTalk video. Use when the user asks for a cover, poster frame, thumbnail, or multiple cover directions; do not use for timeline editing or social-video reframing.
cliptalk-smart-reframe
Creates a subject-aware, time-varying crop track and a review-only social-format preview from an accepted ClipTalk cut. Use for automatic vertical, square, or portrait reframing; do not use for a fixed manual crop or before content editing is accepted.
cliptalk-content-extractor
Locates and assembles source passages matching a semantic request. Use for extracting explanations, topics, quotes, demonstrations, or other specifically described content.
cliptalk-interview-editor
Produces a coherent interview edit by combining speaker discovery, topic selection, dialogue context, cleanup, subtitles, and preview. Use for interviews, podcasts, testimonials, or question-and-answer recordings.
cliptalk-shortform-hook-director
Finds and assembles a reviewable short-form cut with a strong opening hook. Use for Shorts, Reels, social clips, talking-head cutdowns, or requests for a punchier opening.
cliptalk-social-reframe-exporter
Creates a review-only 9:16, 4:5, 1:1, or 16:9 version from an existing accepted ClipTalk cut, then checks the rendered preview. Use only when a cut already exists and the user asks to adapt it for Shorts, Reels, Douyin, Xiaohongshu, WeChat Channels, or square feeds; do not use when the user still needs content found…