Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add shaharsha/claude-skills --skill image-generationgit clone --depth 1 https://github.com/shaharsha/claude-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/shaharsha/claude-skills/image-generation)<a href="https://agentmods.dev/skills/shaharsha/claude-skills/image-generation"><img src="https://agentmods.dev/badge/skills/shaharsha/claude-skills/image-generation/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/shaharsha/claude-skills/image-generation"><img src="https://agentmods.dev/badge/skills/shaharsha/claude-skills/image-generation.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00107 | $0.02838 |
| Opus 5 | $0.00053 | $0.01419 |
| Sonnet 5 | $0.00021 | $0.00568 |
| Haiku 4.5 | $0.00011 | $0.00284 |
Grade A, and why
image-generation scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 166 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Image Generation
Two model families, three jobs.
Models covered
| Model ID | Family | Strengths | Default cost |
|---|---|---|---|
gpt-image-2 |
OpenAI (ChatGPT Images 2.0) | New default. Best-in-class text in any language, near-total prompt adherence, agentic layout reasoning, neutral color accuracy, fast | $0.006 / $0.053 / $0.211 per 1024² at low/medium/high |
gemini-3.1-flash-image-preview |
Google (Nano Banana 2 Flash) | Cheap throwaway exploration; strongest 0.5K preview tier | ~$0.07 / image (1K) |
gemini-3-pro-image-preview |
Google (Nano Banana Pro) | Hyper-realistic portraiture & cinematic lifestyle; character-lock across up to 14 reference images; multi-turn chat editing | ~$0.13 / image (2K), ~$0.24 / image (4K) |
Released 2026-04-21: gpt-image-2 took #1 across every Image Arena category by +242 points over Nano Banana 2 — the largest gap in Arena history. It dethrones gpt-image-1.5 everywhere except native transparent backgrounds, which we now handle via scripts/rembg.sh.
The three jobs of this skill
- Pick the model for the asset and brief. See reference/model-selection.md.
- Write the prompt following the model's house rules. The two providers have OPPOSITE prompt structures — get this wrong and outputs degrade badly:
- OpenAI gpt-image-2 wants labeled segments / line breaks, accepts negative prompts, and rewards explicit constraints. Read reference/openai-gpt-image-2.md.
- Gemini wants narrative paragraphs, negative phrasing actively backfires (rewrite "no people" as "empty street"), aspect ratio goes in
imageConfignot in the prompt text. Read reference/gemini-image.md.
- Run the API via the bundled scripts. See scripts/README.md.
Default model selection — the 30-second rule
Hyper-realistic human face or cinematic portrait?
YES → Gemini Pro 4K.
NO ↓
Need 5+ reference images for character-lock / brand-consistent variants?
YES → Gemini Pro (up to 14 refs with role-assignment).
NO ↓
Throwaway exploration where you'll discard half the outputs?
YES → Gemini Flash 1K or 2K. Promote the keeper to gpt-image-2 high.
NO ↓
Everything else → gpt-image-2 at quality=high.
└── Transparent PNG needed? Generate on flat-white backdrop + scripts/rembg.sh.
See reference/transparent-backgrounds.md.
What ships with it
18 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- examples.md 14 KB
- README.md 5.9 KB
- reference/gemini-image.md 16 KB
- reference/hebrew-rtl.md 7.5 KB
- reference/model-selection.md 11 KB
- reference/openai-gpt-image-2.md 27 KB
- reference/pricing.md 6.3 KB
- reference/transparent-backgrounds.md 6.4 KB
- scripts/gemini-image.sh 6.3 KB runs code
- scripts/openai-image.sh 6.6 KB runs code
- scripts/README.md 6.1 KB
- scripts/rembg.sh 2.4 KB runs code
- templates/hero-image.md 6.5 KB
- templates/icon-set.md 4.5 KB
- templates/logo.md 4.0 KB
- templates/product-shot.md 7.6 KB
- templates/ui-dashboard.md 4.5 KB
- templates/ui-mobile.md 4.2 KB
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 166 lines · 107 tokens per session scan A c6cb445af402
image-generation is a skill published in the GitHub repository shaharsha/claude-skills (5 stars, last pushed 13d ago), licensed MIT. It adds 107 tokens to every session and 2,838 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
html-ppt-zhangzara-retro-zine
A neighborhood zine on the disappearing corner shops — portraits, voices, and what a block loses when they close. Built as a decision-grade story deck for community, local readers.
html-ppt-zhangzara-studio
A photography studio's portfolio-and-rate deck — the signature work, the process, and the packages that win the brief. Built as a decision-grade design craft deck for prospective clients.
motion-frames
A single-frame motion-design composition with looping CSS animations — rotating type ring, animated globe, ticking timer, parallax labels. Renders as a hero video poster you can hand straight to HyperFrames or any keyframe-based exporter. Use when the brief asks for "motion design", "animated hero", "loop", "video…
webgl-halftone-drift
A self-contained WebGL2 hero: a flowing field screened through a rotated halftone dot grid into a duotone print aesthetic; move the cursor to bend the drift.
webgl-holographic-foil
A self-contained WebGL2 hero: thin-film interference over a crushed-foil surface whose palette shifts with the viewing angle; move the cursor to tilt the film.
motion-graphics
A short, design-led motion graphic where motion is the message — kinetic typography, stat count-up, chart/data-viz hit, logo sting / brand lockup, lower-third / callout / social overlay, animated map (highlight regions, connect places, zoom to a location), animated tweet / news-article / headline, webpage / UI…