Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Desko77/cursor-1c-skills --skill transcribegit clone --depth 1 https://github.com/Desko77/cursor-1c-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/desko77/cursor-1c-skills/transcribe)<a href="https://agentmods.dev/skills/desko77/cursor-1c-skills/transcribe"><img src="https://agentmods.dev/badge/skills/desko77/cursor-1c-skills/transcribe/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/desko77/cursor-1c-skills/transcribe"><img src="https://agentmods.dev/badge/skills/desko77/cursor-1c-skills/transcribe.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 1 finding, up to medium
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- medium analysis-evasion · line 1 Suspicious Unicode normalization or mixed-script contentFix: Review the flagged content for security risks. Ensure no credentials, secrets, or sensitive data are exposed.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00196 | $0.07187 |
| Opus 5 | $0.00098 | $0.03594 |
| Sonnet 5 | $0.00039 | $0.01437 |
| Haiku 4.5 | $0.00020 | $0.00719 |
Grade A, and why
transcribe scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
This is a copy
95% identical to transcribe — 15 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
How it starts
The opening of the file, as written. The whole thing — 264 lines — stays where its author put it; the contents beside it link to each section on GitHub.
/transcribe - Транскрибация видео и аудио
Два движка:
- Локальный (default для аудио):
faster-whisper(CUDA) + опц. диаризация. Движок диаризации выбирается сам: без--num-speakers-pyannote community-1(GPU, корректный автодетект числа спикеров, RTF ~0.064); с явным--num-speakers N-sherpa-onnx GPU(pyannote-segmentation-3.0 + eres2net, RTF ~0.24, точное N). Опция--diarize-engine moss- MOSS-Transcribe-Diarize end-to-end: ASR+диаризация одной моделью (без whisper-шага), лучше текст на технических терминах, но ~2x медленнее (RTF ~0.34), требуетvenv-moss(envMOSS_PYTHON). Нет затрат, не уходит наружу. ВИДЕО тоже можно разобрать полностью локально ---engine local(разбор экрана локальной VLM + спикеры по голосу, см. ниже). - Gemini (default для видео и
--analyze-ui): облачный API, ~$0.10/час. Нужен интернет и квота. Стартовая модельgemini-2.5-flash(пин конкретной версии, дешевая); при перегрузке (503/429) переходит наgemini-2.5-flash-lite. Дорогие 3.5/pro сознательно исключены.
Выбор движка по умолчанию
| Тип файла | Движок | Причина |
|---|---|---|
| Аудио (m4a, mp3, wav, ogg, flac, aac, wma) | local | Быстро, бесплатно, диаризация |
| Видео (mp4, mkv, webm, avi, mov) | gemini | Быстро, облако. Приватный вариант - --engine local (см. ниже) |
Видео + --engine local |
local | Разбор экрана БЕЗ облака: whisper + локальная VLM (LM Studio) + спикеры по голосу |
Любой + --analyze-ui |
gemini | Детальный разбор интерфейсов в облаке |
Любой + --engine gemini |
gemini | Явный override на облако |
Аудио + --engine local |
local | Явный override (аудио) |
При 503/429 Gemini-движок сначала сам перебирает пул моделей (см. раздел "Авто-fallback по моделям Gemini"). Если весь пул недоступен и это аудио - можно вручную переключиться на local (--engine local).
Режимы
Локальный (аудио + faster-whisper + опц. pyannote)
Выходные файлы:
<имя> - транскрипция.md- таймкоды + текст<имя> - транскрипция.txt- plain text<имя> - со спикерами.md- реплики с метками[SPEAKER_XX, MM:SS](только при--diarize)
What ships with it
14 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- README.md 14 KB
- scripts/analyze_video_local.py 45 KB runs code
- scripts/diarize_moss.py 8.4 KB runs code
- scripts/diarize_sherpa.py 13 KB runs code
- scripts/glossary.py 12 KB runs code
- scripts/local_backends.py 38 KB runs code
- scripts/setup.py 16 KB runs code
- scripts/speaker_validator.py 38 KB runs code
- scripts/text_stage.py 22 KB runs code
- scripts/transcribe_local.py 37 KB runs code
- scripts/transcribe.py 47 KB runs code
- scripts/verify.py 9.2 KB runs code
- scripts/voiceprints_dedup.py 17 KB runs code
- scripts/voiceprints.py 8.1 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 6d ago First seen · 264 lines · 196 tokens per session scan A f59f14d3a46c
transcribe is a skill published in the GitHub repository Desko77/cursor-1c-skills (55 stars, last pushed 7d ago), licensed MIT. It adds 196 tokens to every session and 7,187 once invoked, about $0.0010 per session on Opus 5. A static security scan graded it A with 0 findings. It is 95% identical to transcribe, differing in 15 lines, and is treated as a copy.
Other skills, from other repositories
mermaid-diagrams
Practical guide for creating human-readable and agent-parseable diagrams using Mermaid. Includes conservative, renderer-compatible templates and when-to-use guidance.
mermaid-render
A tool for turning Mermaid diagram text into PNG, SVG, or PDF files. Mermaid is a text format for describing diagrams such as flows and timelines.
humanize-ai-text
A writing aid for rewriting text produced by language-model agents into a more natural human style. It keeps the original meaning, facts, numbers, and technical terms.
transcribe
A tool for turning video and audio recordings into written text, with optional speaker identification and video analysis.
1c-form-compile
A builder for managed 1C forms, the screens users work with in the 1C business-software platform. It creates the Form.xml file from a JSON description or from an object's metadata.
composing-1c-queries
A guide for writing queries in the 1C:Enterprise query language. It covers how to select, filter, join, group, and summarize data from 1C catalogs, documents, and registers.