Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/moses607/socialforge/visual-idea-generatornpx skills add moses607/socialforge --skill visual-idea-generatorgit clone --depth 1 https://github.com/moses607/socialforgeWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00147 | $0.01269 |
| Opus 5 | $0.00073 | $0.00634 |
| Sonnet 5 | $0.00029 | $0.00254 |
| Haiku 4.5 | $0.00015 | $0.00127 |
Grade A, and why
visual-idea-generator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 57 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Visual Idea Generator
A visual has ~0.4 seconds to win the scroll, so it must communicate ONE idea pre-read: a single subject, one emotion, one promise. Attention is won by contrast and curiosity, not detail — the eye locks onto a human face showing emotion, then reads at most 3 words, then decides. Design from the thumbnail backwards: pick the emotional beat first, compose for a 120px-wide phone preview, and only then write the generation prompt. This skill outputs briefs and prompts you paste into an image tool; it never claims to render pixels itself.
1. Thumbnail psychology (the 5 laws)
- One subject, one focal point. A face, an object, or a before/after — never all three. If you can't name the subject in one word, cut.
- Emotion on a face beats everything. Wide eyes, open mouth, genuine shock/joy. Faces looking AT the viewer or AT the object of interest. This is the single biggest CTR lever.
- Contrast is king. Subject must pop off the background: rim light, blurred/darkened backdrop, or complementary color behind the subject (orange subject on teal, yellow on navy). Avoid busy backgrounds.
- ≤3 words of on-image text, 1 short line, huge and bold (readable at 120px). Text and face must not overlap. Use a punchy word ("WRONG", "$0 → $10K", "DON'T").
- Curiosity gap, not spoiler. Show the stakes, hide the payoff. "I tried it for 30 days" + a shocked face beats "It grew 4x." Open a loop the title/caption closes.
2. Carousel architecture (the 3-act build)
- Cover slide = the hook. Big claim or number + a visual promise of value. Add "swipe" affordance (arrow, "1/7", cut-off next slide). This slide alone decides swipe-through — spend 50% of effort here.
- Value slides (3–8). One idea per slide, consistent template (same margins, font, accent color). Number them. Front-load the best 2 tips — most drop-off happens by slide 3. Use tension → resolution per slide.
- CTA slide. Recap in one line, then ONE ask: follow, save, comment a keyword, or "share this." Never stack CTAs. Restate the handle.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 57 lines · 147 tokens per session scan A 7b678fdb4288
visual-idea-generator is a skill published in the GitHub repository moses607/socialforge (2 stars, last pushed 1mo ago), licensed MIT. It adds 147 tokens to every session and 1,269 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
blog
Full-lifecycle blog engine with 31 sub-skills, 12 content templates, 5-category 100-point scoring, and 5 specialized agents. Routes user requests to the right sub-skill: writing, rewriting, analysis, outlines, audits, schema, charts, images, repurposing, AI citation SEO, FLOW framework prompts, topic-cluster…
blog-notebooklm
Query Google NotebookLM notebooks for source-grounded, citation-backed answers from user-uploaded documents. Manages notebook library, handles Google authentication, and supports smart discovery. Works standalone via /blog notebooklm or internally from blog-write and blog-researcher for source-grounded research…
claude-blog-brain
Scaffold and operate Claude Blog Brain, a source-cited Obsidian brain for blog content creation, optimization, and management dual-optimized for Google rankings (E-E-A-T, the 2026 core updates) and AI citations (GEO/AEO), spanning writing, rewriting and freshness, SERP-informed briefs and outlines, editorial calendars…
blog-google
Google API integration for blog performance: PageSpeed Insights, CrUX Core Web Vitals with 25-week history, Search Console performance, URL Inspection, Indexing API, GA4 organic traffic, NLP entity analysis for E-E-A-T, YouTube video search for embedding, and Google Ads Keyword Planner. Progressive feature…
banana
Direct, generate, edit, compare, and review visual assets with current Google Gemini image models. Use for image creation, image editing, reference-based consistency, product and character visuals, text-bearing graphics, grounded diagrams, video-derived images, and multi-model image portfolios.
narrator-ai-cli
AI 电影/短剧解说视频自动生成(AI 解说大师 CLI Skill)。当用户需要创建电影解说视频、短剧解说、影视二创、AI 配音旁白视频、film commentary、video narration、drama dubbing、movie narration 时触发。内置电影素材库、BGM、多语种配音、解说模板。通过 narrator-ai-cli 命令行实现:搜片→选模板→选 BGM→选配音→生成文案→合成视频的全流程自动化。CLI client for Narrator AI video narration API.