describe-frame

describe-frame is an agent for Claude Code from pamelafox/presentation-skills. It costs 2 tokens per session (457 once invoked), scanned A, original, MIT.

An agent that describes what appears in a single frame from a presentation video or compares two frames for speaker-face quality. It can note slide content, visible activity, and whether speakers are facing the camera with their eyes and mouth open or closed.

In plain words
What is it for?
Use it to caption presentation-video frames, detect meaningful slide changes, and compare two speaker frames.
Why use it?
It provides consistent text descriptions for changing video frames and can identify when a frame is effectively unchanged. It also helps select or compare frames where a speaker's face is clearly visible.

Agent for Claude Code

Written for Claude Code: user-invocable in frontmatter.

Good fit Use it to caption presentation-video frames, detect meaningful slide changes, and compare…

Compare 6 agents from other repositories ↓
Install with agentmods
npx agentmods add agents/pamelafox/presentation-skills/describe-frame
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Clone the repo
git clone --depth 1 https://github.com/pamelafox/presentation-skills

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for describe-frame

README.md
[![agentmods](https://agentmods.dev/badge/agents/pamelafox/presentation-skills/describe-frame.svg)](https://agentmods.dev/agents/pamelafox/presentation-skills/describe-frame)
Your own site
<a href="https://agentmods.dev/agents/pamelafox/presentation-skills/describe-frame"><img src="https://agentmods.dev/badge/agents/pamelafox/presentation-skills/describe-frame.svg" alt="Measured on agentmods" height="20"></a>
Per session 2 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 457 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00002 $0.00457
Opus 5 $0.00001 $0.00229
Sonnet 5 $0.00000 $0.00091
Haiku 4.5 $0.00000 $0.00046

Measured 7d ago against content hash 40e196c02581, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-06, from the pricing page.

Security

Grade A, and why

describe-frame scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.github/agents/describe-frame.md · 38 lines

What it actually says

You describe a single frame captured from a presentation video, or compare two frames for speaker face quality.

You will be given one of two tasks:

Task A: Describe a frame

  • The path to the current frame image to view.
  • Optionally, the path to the previous frame image to view.
  • Optionally, the previous frame's description as text.

Task B: Compare speaker faces

  • Paths to two frame images to compare.
  • Instructions to evaluate speaker face quality (eyes open, mouth open, facing camera).

Instructions

For Task A (description):

  1. View the current frame image.
  2. If a previous frame image path is provided, view it as well.
  3. Compare the two frames. If they show essentially the same content (same slide, no meaningful change), respond with ONLY: (same as previous)
  4. Otherwise, write a one-sentence description of the current frame focusing on visible slide content, diagrams, code, or speaker activity.
  5. If any speaker faces are visible (webcam feeds, on-camera presenter), append face attributes in parentheses at the end of the description. For each visible speaker, note: mouth open/closed, eyes open/closed, facing camera or turned away. Example: (Speaker A: mouth open, eyes open, facing camera; Speaker B: mouth closed, eyes open, looking down)
  6. Respond with ONLY the description text — no JSON, no markdown formatting, no extra commentary.

For Task B (face comparison):

  1. View both frame images.
  2. For each frame, check all visible speaker faces for:
    • Mouth: open (mid-speech) or closed
    • Eyes: open or closed/mid-blink
    • Orientation: facing camera or turned away/looking down
  3. Reply with the requested format (e.g., "A BETTER", "B BETTER", "EQUAL", or "NO SPEAKERS") followed by a brief reason.
  4. When checking if a specific speaker's mouth is open, focus on that speaker only.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 7d ago First seen · 38 lines · 2 tokens per session scan A 40e196c02581

Subscribe to this mod's changes

describe-frame is an agent published in the GitHub repository pamelafox/presentation-skills (117 stars, last pushed 3d ago), licensed MIT. It adds 2 tokens to every session and 457 once invoked, about $0.0000 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

pixel-art-animation-reviewer

Independent reviewer of pixel-art ANIMATION quality (loop seamlessness, motion physics, multi-component motion, frame timing, period selection, particle determinism). One of four specialized review roles in the pixel-art-quality-board orchestrator. Use when the user asks to "check animation timing", "verify loop…

AnastasiyaW/codex-claude-code-config · 140 tokens

proposal-writer

Specialized agent for generating professional, branded proposals using a presentation-generation tool. Creates polished presentations and documents for sales opportunities from your project and CRM context.

Zeekeey-jpeg/LeRoy-HQ · 34 tokens

cover-artist

Generate book cover art prompts from story content. Produces optimized prompts for image generation models (GPT Image, Gemini, FLUX, etc.) that conform to Kindle dimensions.

howells/fiction · 38 tokens

ollama-vision

Use this agent to analyze images, screenshots, UI mockups, diagrams, or any visual content. Delegates vision analysis to a local Qwen2.5-VL model. Use when the user wants to describe, debug, or extract information from an image file.

PratikHotchandani22/claude-ollama-agents · 59 tokens

forge-modeler

Headless 3D geometry specialist for the Forge suite. Builds, repairs, and validates polygon meshes, parametric CAD (CadQuery/Build123d/OpenSCAD), and procedural geometry (Geometry Nodes, SDF, L-systems) via Python — no GUI. Use for mesh construction, parametric modeling, procedural generation, topology/retopo/LOD…

luminary19/atelier · 126 tokens

gds-agent-game-designer

Game designer for creative vision, GDD creation, and narrative design. Use when the user asks to talk to Samus Shepard or requests the Game Designer.

PabloLION/bmad-plugin · 38 tokens