video-layer-skill

video-layer-skill is a skill for Claude Code, Codex from Alexander-Kz/video-layer-skill. It costs 183 tokens per session (8,187 once invoked), scanned A, original, MIT.

A workflow for producing whiteboard-style explainer video layers from a voiceover MP3, with scenes planned and synchronized to the narration.

In plain words
What is it for?
Use it to transcribe narration, plan rapid scene changes, write image prompts, and coordinate the agents and scripts needed to build the video layers.
Why use it?
It organizes the work of turning spoken content into a sequence of visual scenes for a faceless explainer video.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one. Also seen: reads .claude/ paths; mentions CLAUDE.md; mentions subagents.

Good fit Use it to transcribe narration, plan rapid scene changes, write image prompts, and coordinate the agents and scripts needed to build the video layers.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/alexander-kz/video-layer-skill/video-layer-skill
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add Alexander-Kz/video-layer-skill --skill video-layer-skill
Clone the repo
git clone --depth 1 https://github.com/Alexander-Kz/video-layer-skill

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for video-layer-skill

README.md
[![agentmods](https://agentmods.dev/badge/skills/alexander-kz/video-layer-skill/video-layer-skill/github.svg)](https://agentmods.dev/skills/alexander-kz/video-layer-skill/video-layer-skill)
Your own site
<a href="https://agentmods.dev/skills/alexander-kz/video-layer-skill/video-layer-skill"><img src="https://agentmods.dev/badge/skills/alexander-kz/video-layer-skill/video-layer-skill/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for video-layer-skill

Your own site · 80×15
<a href="https://agentmods.dev/skills/alexander-kz/video-layer-skill/video-layer-skill"><img src="https://agentmods.dev/badge/skills/alexander-kz/video-layer-skill/video-layer-skill.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 183 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 8,187 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00183 $0.08187
Opus 5 $0.00092 $0.04093
Sonnet 5 $0.00037 $0.01637
Haiku 4.5 $0.00018 $0.00819

Measured 9d ago against content hash 0b8e2e5391bf, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-09, from the pricing page.

Security

Grade A, and why

video-layer-skill scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

SKILL.md · 780 lines

How it starts

The opening of the file, as written. The whole thing — 780 lines — stays where its author put it; the contents beside it link to each section on GitHub.

video-layer-skill — Director Orchestration

You are the Director of a multi-agent pipeline that produces whiteboard-style explainer videos from a voiceover audio file. Your job: conduct a short interview with the user, then orchestrate Planner / Writer / Reviewer agents through a strict sequence of phases. You do NOT do heavy work yourself — you spawn agents, run scripts, and coordinate.

Your intelligence is free (Claude Max). What is expensive: image generation via Replicate. Always show cost estimates and get user approval before spending money.


Working language

Communicate with the user in their language (Russian by default for this user). All system prompts, model prompts, and generated content are English (image models require English).


Mission (every agent must obey)

Produce a video where:

  1. Visual style is classic whiteboard hand-drawn animation — thick black marker lines, flat solid colors from a strict palette, white "paper" background, optional stick figures with closed mouths.
  2. Pacing is rapid and deliberately varied — image-change frequency targets:
    • Hook (0–10 s): ~5–6 image changes total. Median image hold ~1.5 s (range 1.3–1.7 s). One sub-1 s cut is fine; do NOT chain sub-1 s cuts. Hook punch comes from the narration line, not from cut speed.
    • Body (after ~10 s): target ~22 cuts/min. Median image hold ~2.6 s. Most cuts fall in the 1.7–3.7 s "walking pace" band.
    • Sustained holds (3–7.5 s) are required, not optional. Plan 2–3 holds of 5–7 s per minute on landmark beats: key reveal, emotional peak, mid-sentence pause, dense infographic / multi-element scene that needs reading time, single evocative image carrying a whole sentence.
    • Short bursts (<1 s, max 3 in a row) allowed for enumerations, climactic reveals, comedic beats, energy spikes.
    • Distribution target across whole video: ~25–30% under 1.5 s, ~30–35% at 1.5–3 s, ~25–30% at 3–5 s, ~10–15% at 5–7.5 s.
    • Hard cap: 7.5 s per scene. Hard min: 0.5 s.
  3. Cuts are hard cuts (no Ken Burns, no transitions, no fades).
  4. Every image is clean — no text/captions/watermarks unless narratively required (named numbers, years, key terms — Nano Banana 2 renders text well).
  5. Sequence chains (≤4 frames showing progression in one location) are generated strictly serially, each frame using the previous as reference.

Read the full file on GitHub · 780 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 9d ago First seen · 780 lines · 183 tokens per session scan A 0b8e2e5391bf

Subscribe to this mod's changes

video-layer-skill is a skill published in the GitHub repository Alexander-Kz/video-layer-skill (5 stars, last pushed 3mo ago), licensed MIT. It adds 183 tokens to every session and 8,187 once invoked, about $0.0009 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

shogun-screenshot

A screenshot tool for getting images from a computer or web page and then cropping, resizing, or masking sensitive information. Playwright is a browser-automation tool used here to capture web pages.

yohey-w/multi-agent-shogun · 149 tokens

guizang-social-card-skill

Generate Guizang-style social card image sets, Live Photo motion cards, material-first Live Photo puzzle layouts, triple Live Photo collages, long-video-to-Live-Photo treatments, and WeChat official account cover pairs from articles, scripts, screenshots, product notes, subtitles, photos, or user-supplied videos. Use…

op7418/guizang-social-card-skill · 170 tokens

html-effectiveness-diagram

A guide for creating illustrations and diagrams directly in a web page with inline SVG, a built-in format for scalable vector graphics. It covers document illustrations, flowcharts, state machines, and deployment pipelines.

Azhi-ss/html-effectiveness · 95 tokens

html-effectiveness-slides

A format for small, single-file HTML slide decks that advance with the keyboard. HTML is the language used to structure web pages.

Azhi-ss/html-effectiveness · 82 tokens

canvas-design

Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.

shahtuyakov/claude-setup · 59 tokens

ascii-video

ASCII video: convert video/audio to colored ASCII MP4/GIF.

NousResearch/hermes-agent · 17 tokens