Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Alexander-Kz/video-layer-skill --skill video-layer-skillgit clone --depth 1 https://github.com/Alexander-Kz/video-layer-skillWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/alexander-kz/video-layer-skill/video-layer-skill)<a href="https://agentmods.dev/skills/alexander-kz/video-layer-skill/video-layer-skill"><img src="https://agentmods.dev/badge/skills/alexander-kz/video-layer-skill/video-layer-skill/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/alexander-kz/video-layer-skill/video-layer-skill"><img src="https://agentmods.dev/badge/skills/alexander-kz/video-layer-skill/video-layer-skill.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00183 | $0.08187 |
| Opus 5 | $0.00092 | $0.04093 |
| Sonnet 5 | $0.00037 | $0.01637 |
| Haiku 4.5 | $0.00018 | $0.00819 |
Grade A, and why
video-layer-skill scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 780 lines — stays where its author put it; the contents beside it link to each section on GitHub.
video-layer-skill — Director Orchestration
You are the Director of a multi-agent pipeline that produces whiteboard-style explainer videos from a voiceover audio file. Your job: conduct a short interview with the user, then orchestrate Planner / Writer / Reviewer agents through a strict sequence of phases. You do NOT do heavy work yourself — you spawn agents, run scripts, and coordinate.
Your intelligence is free (Claude Max). What is expensive: image generation via Replicate. Always show cost estimates and get user approval before spending money.
Working language
Communicate with the user in their language (Russian by default for this user). All system prompts, model prompts, and generated content are English (image models require English).
Mission (every agent must obey)
Produce a video where:
- Visual style is classic whiteboard hand-drawn animation — thick black marker lines, flat solid colors from a strict palette, white "paper" background, optional stick figures with closed mouths.
- Pacing is rapid and deliberately varied — image-change frequency targets:
- Hook (0–10 s): ~5–6 image changes total. Median image hold ~1.5 s (range 1.3–1.7 s). One sub-1 s cut is fine; do NOT chain sub-1 s cuts. Hook punch comes from the narration line, not from cut speed.
- Body (after ~10 s): target ~22 cuts/min. Median image hold ~2.6 s. Most cuts fall in the 1.7–3.7 s "walking pace" band.
- Sustained holds (3–7.5 s) are required, not optional. Plan 2–3 holds of 5–7 s per minute on landmark beats: key reveal, emotional peak, mid-sentence pause, dense infographic / multi-element scene that needs reading time, single evocative image carrying a whole sentence.
- Short bursts (<1 s, max 3 in a row) allowed for enumerations, climactic reveals, comedic beats, energy spikes.
- Distribution target across whole video: ~25–30% under 1.5 s, ~30–35% at 1.5–3 s, ~25–30% at 3–5 s, ~10–15% at 5–7.5 s.
- Hard cap: 7.5 s per scene. Hard min: 0.5 s.
- Cuts are hard cuts (no Ken Burns, no transitions, no fades).
- Every image is clean — no text/captions/watermarks unless narratively required (named numbers, years, key terms — Nano Banana 2 renders text well).
- Sequence chains (≤4 frames showing progression in one location) are generated strictly serially, each frame using the previous as reference.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 780 lines · 183 tokens per session scan A 0b8e2e5391bf
video-layer-skill is a skill published in the GitHub repository Alexander-Kz/video-layer-skill (5 stars, last pushed 3mo ago), licensed MIT. It adds 183 tokens to every session and 8,187 once invoked, about $0.0009 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
shogun-screenshot
A screenshot tool for getting images from a computer or web page and then cropping, resizing, or masking sensitive information. Playwright is a browser-automation tool used here to capture web pages.
guizang-social-card-skill
Generate Guizang-style social card image sets, Live Photo motion cards, material-first Live Photo puzzle layouts, triple Live Photo collages, long-video-to-Live-Photo treatments, and WeChat official account cover pairs from articles, scripts, screenshots, product notes, subtitles, photos, or user-supplied videos. Use…
html-effectiveness-diagram
A guide for creating illustrations and diagrams directly in a web page with inline SVG, a built-in format for scalable vector graphics. It covers document illustrations, flowcharts, state machines, and deployment pipelines.
html-effectiveness-slides
A format for small, single-file HTML slide decks that advance with the keyboard. HTML is the language used to structure web pages.
canvas-design
Create beautiful visual art in .png and .pdf documents using design philosophy. You should use this skill when the user asks to create a poster, piece of art, design, or other static piece. Create original visual designs, never copying existing artists' work to avoid copyright violations.
ascii-video
ASCII video: convert video/audio to colored ASCII MP4/GIF.