Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add naveedharri/benai-skills --skill benai-skill-creator-skillgit clone --depth 1 https://github.com/naveedharri/benai-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/naveedharri/benai-skills/benai-skill-creator-skill)<a href="https://agentmods.dev/skills/naveedharri/benai-skills/benai-skill-creator-skill"><img src="https://agentmods.dev/badge/skills/naveedharri/benai-skills/benai-skill-creator-skill/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/naveedharri/benai-skills/benai-skill-creator-skill"><img src="https://agentmods.dev/badge/skills/naveedharri/benai-skills/benai-skill-creator-skill.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 1 finding, up to medium
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- medium Excessive Agency · line 41 Skill enables autonomous high-impact decisions without human-in-the-loop verification. Critical operations (destructive commands, financial transactions, data deletion) should require explicit user confirmation.Fix: Add human-in-the-loop confirmation for destructive, irreversible, or high-impact operations. Never auto-execute commands that modify files, send data, or alter system state.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00190 | $0.01114 |
| Opus 5 | $0.00095 | $0.00557 |
| Sonnet 5 | $0.00038 | $0.00223 |
| Haiku 4.5 | $0.00019 | $0.00111 |
Grade A, and why
benai-skill-creator-skill scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 58 lines — stays where its author put it; the contents beside it link to each section on GitHub.
BenAI Skill Creator
Most skills are built wrong: someone describes a process from memory before running it, and gets the cleaned-up version that drops the judgment and the exceptions. This does the opposite. You do the task once in a chat, then this turns what actually happened into a small, reliable skill.
Two rules carry the whole skill: keep it modular (one goal, one job) and remember skills are never finished (ship a small working one, then improve it by using it). The reasoning is in references/scope-and-mindset.md.
Route first
Read the situation and pick a branch. Say which branch you picked and why, then proceed.
| Situation | Branch | Do this |
|---|---|---|
| The user just finished a real task in this chat | Build | Steps 1 to 5 below |
| The user only has an idea, nothing has been done yet | Plan | Hand off to the process-interviewer skill (interview first, then it builds). Stop here. |
| The user points at an existing skill to fix or shrink | Improve | Skip step 2 (Extract). Run steps 1, 3, 4, 5 against the existing skill. |
If unsure, ask one question: "Did we already do this task in this chat, or is it still just an idea?"
The build flow
Track progress out loud:
Task Progress:
- [ ] 1. Scope: one goal, right size, split or not
- [ ] 2. Extract: pull the real process out of the chat
- [ ] 3. Structure: write a small SKILL.md + reference files
- [ ] 4. Prune: cut everything that does not change behavior
- [ ] 5. Eval: test it works, add the self-improvement rule
1. Scope
Decide the one job and whether it should be one skill or several. Read references/scope-and-mindset.md. The cutoff line: a skill is one task doable in a single chat session (roughly 1 to 3 prompts). If the task is bigger, split it into a chain of small skills.
2. Extract
Reverse-engineer what actually happened. Do not ask the user to re-describe their process from memory. Walk back through this conversation and pull out the real steps, the judgment calls, and the reference material. Read references/extract-from-task.md for the technique, then play the process back to the user and let them correct it.
What ships with it
12 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- references/build-methods.md 1.3 KB
- references/connectors-and-mcp.md 990 B
- references/evals-and-improvement.md 1.7 KB
- references/example-churn-recovery.md 2.4 KB
- references/extract-from-task.md 1.8 KB
- references/prune.md 1.4 KB
- references/scope-and-mindset.md 2.2 KB
- references/structure.md 4.7 KB
- templates/_reference.md.tmpl 564 B
- templates/_SKILL.md.tmpl 1.3 KB
- templates/eval-prompt.md 1.4 KB
- templates/self-improvement-rule.md 1.0 KB
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 58 lines · 190 tokens per session scan A 3a59f15a213d
benai-skill-creator-skill is a skill published in the GitHub repository naveedharri/benai-skills (61 stars, last pushed today), licensed MIT. It adds 190 tokens to every session and 1,114 once invoked, about $0.0010 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
project
A single starting point for setting up an AI-assisted project in Claude Cowork. It asks questions about the work, reviews available add-ons, and creates project instructions, custom agents, and connected workflows.
doc-html-slide
A renderer that turns presentation content into a single HTML slide deck that opens directly in a browser. It creates a 16:9 slide sequence with navigation, fullscreen viewing, printing to PDF, and speaker-note controls.
cs-channel-message
A channel-message writing tool for search ads, advertising, customer relationship messages, and app notifications. It uses the NCM sequence—Need, Channel, Moment, Message, and CTA—to adapt wording to where and when customers see it.
design-sync-upload
An uploader for design-system files such as DESIGN.md, tokens, logos, fonts, and images into Claude Design. It can either use an authenticated connection or prepare a folder and guide for manual upload.
design-tokens-transformer
A converter for design tokens, which are named values for colors, fonts, spacing, borders, shadows, and motion. It translates one shared token source into CSS variables and Tailwind or shadcn-style theme files, and can convert them back for checking.
media-higgsfield-explainer
A Higgsfield workflow for making non-photorealistic narrated explainer videos. It pairs each narration line with a 10-second animated clip and joins the clips into one finished video.