Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add OpenLAIR/OpenSkill --skill evo-speaker-diarization-subtitlesgit clone --depth 1 https://github.com/OpenLAIR/OpenSkillWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/openlair/openskill/evo-speaker-diarization-subtitles)<a href="https://agentmods.dev/skills/openlair/openskill/evo-speaker-diarization-subtitles"><img src="https://agentmods.dev/badge/skills/openlair/openskill/evo-speaker-diarization-subtitles/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/openlair/openskill/evo-speaker-diarization-subtitles"><img src="https://agentmods.dev/badge/skills/openlair/openskill/evo-speaker-diarization-subtitles.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00093 | $0.00635 |
| Opus 5 | $0.00046 | $0.00318 |
| Sonnet 5 | $0.00019 | $0.00127 |
| Haiku 4.5 | $0.00009 | $0.00064 |
Grade A, and why
evo-speaker-diarization-subtitles scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
evo-speaker-diarization-subtitles
End-to-end speaker diarization and subtitle generation pipeline for video/audio files.
Pipeline
- Audio extraction: ffmpeg → 16kHz mono 16-bit PCM WAV
- Voice Activity Detection: silero-vad (with fallback to torch.hub if direct import unavailable)
- Speaker embedding extraction: SpeechBrain ECAPA-TDNN (
spkrec-ecapa-voxceleb) - Speaker clustering: Agglomerative clustering with cosine distance + silhouette-based speaker count selection
- Transcription: OpenAI Whisper with word-level timestamps
- Alignment: Word-level IoU overlap mapping from Whisper words to diarized speaker segments
- Output generation: RTTM (pyannote.core or manual), ASS subtitles, JSON report
Usage
import sys
sys.path.insert(0, '/app/environment/skills/evo-speaker-diarization-subtitles/scripts')
from utils import (
extract_audio, get_audio_duration, run_vad,
merge_close_segments, extract_speaker_embeddings,
cluster_speakers, run_whisper_transcription,
align_transcription_with_speakers,
write_rttm, write_ass, write_report,
format_time_ass
)
Dependencies
ffmpeg, torch, torchaudio, speechbrain, openai-whisper, silero-vad, scikit-learn, numpy, soundfile, scipy, pyannote.core
Key Domain Notes
- ECAPA-TDNN embeddings are trained with angular margin loss → cosine distance is the correct metric for clustering.
- Silero VAD v6.2.0 uses
from silero_vad import load_silero_vad, get_speech_timestamps(not torch.hub). Fallback to torch.hub is provided for older versions. - Whisper
word_timestamps=Trueenables word-level alignment critical for accurate speaker attribution. - ASS time format uses centiseconds:
H:MM:SS.cc— careful rollover handling is required. - RTTM format:
SPEAKER <file_id> 1 <start> <duration> <NA> <NA> <speaker_label> <NA> <NA> - Speaker labels in RTTM use
spkNNformat; in ASS subtitles useSPEAKER_NNformat. - Segments shorter than ~0.5s yield unreliable embeddings; minimum 0.15s is enforced, 0.5s preferred.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 45 lines · 93 tokens per session scan A e90e4c33cdb3
evo-speaker-diarization-subtitles is a skill published in the GitHub repository OpenLAIR/OpenSkill (90 stars, last pushed 2d ago), licensed Apache-2.0. It adds 93 tokens to every session and 635 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-11.
Other skills, from other repositories
flux-analyzer
Analyse FBA flux distributions to extract biological insights. Covers gene essentiality, phenotypic phase planes, flux sampling, pathway-level aggregation, secretion product prediction, and production of publication- quality figures.
experimental-design
Best practices for designing reproducible ML experiments. Use when planning ablations, baselines, or controlled experiments.
prolong
Recover and use durable coding-session history from PRO-LONG's local append-only log. Use on long-running coding tasks, after context compaction or session resume, when reconstructing prior decisions or tool results, or before repeating work that may already have been attempted.
resolving-merge-conflicts
Use when a git merge or rebase reports conflicts and the operation is in progress.
security-review
Use when reviewing code for vulnerabilities, checking diffs for injection/XSS/auth/crypto issues. Invokes on code changes (.go/.py/.js/.ts/.java/.rs/.php/.rb), not docs-only diffs.
script-exec-blocked
Sandbox approval policy blocks executecode and python3 -c; use readfile/writefile + manual transforms instead of retrying both runners.