Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add calesthio/generative-media-skills --skill performance-directiongit clone --depth 1 https://github.com/calesthio/generative-media-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/calesthio/generative-media-skills/performance-direction)<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/performance-direction"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/performance-direction/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/calesthio/generative-media-skills/performance-direction"><img src="https://agentmods.dev/badge/skills/calesthio/generative-media-skills/performance-direction.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00101 | $0.04887 |
| Opus 5 | $0.00051 | $0.02443 |
| Sonnet 5 | $0.00020 | $0.00977 |
| Haiku 4.5 | $0.00010 | $0.00489 |
Grade A, and why
performance-direction scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 349 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Performance direction for generated media
Use performance direction to make a generated person or character appear to want something, react to pressure, and communicate through voice, face, gaze, and body. Do not ask a model for a generic emotion and hope it supplies acting. Give playable circumstances, beat-level actions, physical behavior, vocal behavior, continuity constraints, and a way to judge the take.
Evidence stance
Documented facts used by this skill:
- Stanislavski-derived actor training commonly analyzes a role through given circumstances, tasks/objectives, actions, beats, obstacles, and physical action; for AI use, translate these into concise, playable instructions rather than academic labels. Source cross-check: Stanislavski system overview and NYFA summary of Stanislavski questions.
- Lip-sync systems often map speech sound to visemes, not one mouth shape per phoneme. Microsoft documents visemes as visual phoneme descriptions and notes many phonemes can share one viseme; Meta's Oculus Lipsync documentation describes interpolation between visemes over time. Verified 2026-07-10: Microsoft viseme docs, Meta viseme reference.
- Professional dubbing guidance treats timing, natural language, audio level, and physical mouth/gesture sync as production constraints. Verified 2026-07-10: Netflix films/series dubbing guidelines, Netflix nonfiction simil-sync guidance.
- Current avatar tools expose different levels of performance control: some require clean training footage and consent capture, some accept a script and avatar selection, and some offer sentiment/emotional-state controls. Verified 2026-07-10: Synthesia Studio Avatar requirements, D-ID quickstart, HeyGen Digital Twin docs.
- Current AI video prompting can include character references, longer clips, extension, and batch workflows in some providers; treat capability limits as provider- and date-specific. Verified 2026-07-10: OpenAI Sora 2 prompting guide.
- Performer likeness, voice, and digital replica use raise consent, compensation, use-limit, storage, and disclosure duties. Verified 2026-07-10: SAG-AFTRA AI resource page, SAG-AFTRA Digital Replicas PDF, NAVA synthetic voice guidance.
- Provenance/disclosure can be supported by metadata standards such as C2PA Content Credentials, but provenance support varies by platform and workflow. Verified 2026-07-10: C2PA.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 349 lines · 101 tokens per session scan A 7a81000d4066
performance-direction is a skill published in the GitHub repository calesthio/generative-media-skills (170 stars, last pushed 1mo ago), licensed MIT. It adds 101 tokens to every session and 4,887 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
feature-demo-recording
Record a demo video of a web feature from a real browser. Two modes -- a NARRATED film where measured voiceover drives the timeline (designed slides, subtitles, punch-in camera, rendered from an HTML timeline), and a SILENT evidence clip for a PR or a QA pass. Use when the user asks to record a video, demo, or screen…
image-authoring
Author images and diagrams as code — SVG, Pillow, Excalidraw, mermaid. Load when asked to draw, illustrate, or make an image, icon, logo, poster, or diagram.
bento-slides
Create and edit Bento presentations — self-contained .bento.html decks whose document is JSON. Use whenever the user wants a slide deck or presentation: from scratch, from source material, or by improving an existing file.
pptx-maker
Generate or restyle a PowerPoint deck. Use when the user wants to create or edit a .pptx presentation, build slides from text or a URL, or design a reusable slide style.
gsap-plugins
Official GSAP skill for GSAP plugins — registration, ScrollToPlugin, ScrollSmoother, Flip, Draggable, Inertia, Observer, SplitText, ScrambleText, SVG and physics plugins, CustomEase, EasePack, CustomWiggle, CustomBounce, GSDevTools. Use when the user asks about a GSAP plugin, scroll-to, flip animations, draggable, SVG…
seedance-2-0
Generate cinematic clips with ByteDance Seedance 2.0 — the preferred premium video model in OpenMontage when a paid gateway is configured. Use when: (1) producing trailers, teasers, hype edits, or premium cinematic clips, (2) needing native synchronized audio (speech, SFX, ambience) in a single pass, (3) needing…