Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/hex/claude-image-generation/image-generationnpx skills add hex/claude-image-generation --skill image-generationgit clone --depth 1 https://github.com/hex/claude-image-generationWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/hex/claude-image-generation/image-generation)<a href="https://agentmods.dev/skills/hex/claude-image-generation/image-generation"><img src="https://agentmods.dev/badge/skills/hex/claude-image-generation/image-generation.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00127 | $0.03651 |
| Opus 5 | $0.00063 | $0.01826 |
| Sonnet 5 | $0.00025 | $0.00730 |
| Haiku 4.5 | $0.00013 | $0.00365 |
Grade A, and why
image-generation scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 239 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Image Generation with Gemini, OpenAI, xAI, and OpenRouter
Generate and edit images using Google Gemini, OpenAI GPT Image 2, xAI Grok Image, and OpenRouter APIs via shell scripts.
Available Providers
Google Gemini
- Model:
gemini-3-pro-image(default, "Nano Banana Pro"). Alt:gemini-3.1-flash-image(Flash, 14 ratios),gemini-3.1-flash-lite-image(cheapest). The-previewIDs are past their shutdown date. - Strengths: Premium quality, up to 4K output, thinking mode, Google Search grounding, multi-turn editing with up to 14 reference images
- Aspect ratios: 10 on Pro (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 4:5, 5:4, 21:9); Flash adds 4 extreme ratios (1:4, 4:1, 1:8, 8:1)
- Resolution:
--image-sizetakes1K,2K,4Kon both Pro and Flash; Flash additionally supports512(UPPERCASE required) - Env var:
GEMINI_API_KEY
OpenAI GPT Image 2
- Model:
gpt-image-2(default, snapshotgpt-image-2-2026-04-21);gpt-image-1.5available as previous flagship via--model - Strengths: Superior text rendering, transparent backgrounds, up to 16 input images for editing, quality tiers
- Sizes: 1024x1024, 1536x1024 (landscape), 1024x1536 (portrait)
- Quality: low (fast/cheap), medium, high (best fidelity)
- Env var:
OPENAI_API_KEY
xAI Grok Image
- Model:
grok-imagine-image-2.0(default, flagship since 2026-08-07),grok-imagine-image-quality(May 2026 quality mode;grok-imagine-image-proredirects here),grok-imagine-image(standard, 300 RPM) - Strengths: Prompt revision by chat model, flat per-image pricing, diverse style range, many aspect ratios
- Aspect ratios: 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3, 2:1, 1:2, 19.5:9, 9:19.5, 20:9, 9:20, 21:9, 5:2, auto
- Resolution:
--resolutiontakes1k,2k(LOWERCASE required, opposite of Gemini) - Quality:
--qualitytakeslow,medium,autoongrok-imagine-image-2.0; unset means auto (low for generation, medium for edits), billed as served - Editing: Dedicated
/v1/images/editsendpoint; up to 5 input images passed as data URIs in animagesarray - Env var:
XAI_API_KEYorGROK_API_KEY
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- today Changed · +49 lines · +14 tokens per session 9431508c8909
- 4d ago First seen · 190 lines · 113 tokens per session scan A 0c406971dd66
image-generation is a skill published in the GitHub repository hex/claude-image-generation (8 stars, last pushed yesterday), licensed MIT. It adds 127 tokens to every session and 3,651 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
crossgen-artist
Use CrossGen to plan and execute image generation, editing, inpainting, model selection, job monitoring, Gallery inspection, and asset export through MCP or its JSON CLI.
ima2
Use the ima2-gen CLI/server to generate, edit, inspect, and manage local AI image generation jobs.
ima2-front
Frontend implementation skill for ima2 users. Use for any frontend, web UI, or visual implementation work — building, styling, or redesigning pages/components, responsive layouts, motion, component architecture, and production-surface polish. Pairs with ima2-uiux: load it first when design direction is vague; this…
deep-execution
Executes agent-enhanced council queries by spawning parallel Claude subagents that each query a provider, evaluate response quality, ask follow-up questions, and return structured insights with confidence ratings and blind spot analysis. Invoked when the --agents flag is used or when complex architectural decisions…
videoagent-image-studio
Tired of juggling 8 API keys? This skill gives you one-command access to Midjourney, Flux, Ideogram, and more, with zero setup. Use when you want to generate any image without worrying about API keys.
video-perception
Use when the user mentions a video file (.mp4, .mov, .avi, .mkv, .webm), a YouTube URL, asks to watch/analyze/review a video, or references video content in conversation.