skill-enhancer

A guide for auditing and improving existing agent skills against defined quality rules. TDD means test-driven development: checking expected behavior through tests while building or changing something.

In plain words
What is it for?
Use it to analyze a skill's gaps, improve its wording and structure, audit its execution and safety rules, and address validator findings.
Why use it?
It identifies missing instructions, weak descriptions, unsafe practices, and validation gaps before a skill is treated as complete.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/matrixfounder/agentic-development/skill-enhancer
Any agent
npx skills add MatrixFounder/Agentic-development --skill skill-enhancer
Clone the repo
git clone --depth 1 https://github.com/MatrixFounder/Agentic-development

Made for: Claude Code, Codex.

Per session 25 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,292 The whole file, excluding the scripts and references it only reads on demand.
Security scan B 2 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00025 $0.02292
Opus 5 $0.00013 $0.01146
Sonnet 5 $0.00005 $0.00458
Haiku 4.5 $0.00003 $0.00229

Measured 2d ago against content hash f2417392b5a8, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade B, and why

skill-enhancer scanned grade B with 2 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

The scan reads SKILL.md. This mod also ships 2 executable files (scripts/analyze_gaps.py, scripts/skill_utils.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Downloads and executes remote codemediumSupply chain

curl | sh runs whatever the server returns today, which is not necessarily what it returned when this was reviewed.

- **Security Remediation**: Fix vulnerabilities flagged by `skill-validator` (e.g., `curl | bash`, secrets, weak permissions).

Downgraded: this mod is about security review, or the phrase is quoted, so it is likely naming the pattern rather than instructing it.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

- **Security Remediation**: Fix vulnerabilities flagged by `skill-validator` (e.g., `curl | bash`, secrets, weak permissions).
.agent/skills/skill-enhancer/SKILL.md · 150 lines

How it starts

The opening of the file, as written. The whole thing — 150 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Skill Enhancer

Purpose: This meta-skill analyzes other skills for compliance with TDD, CSO, and Script-First standards, guiding the agent through upgrades.

1. Red Flags (Anti-Rationalization)

STOP and READ THIS if you are thinking:

  • "I'll just add the sections blindly" -> WRONG. You must understand why the skill fails before fixing it.
  • "The description is close enough" -> WRONG. It must start with "Use when".
  • "Examples are optional" -> WRONG. "Rich Skills" mandate examples.
  • "It's just a small 20-line example" -> WRONG. Inline blocks > 12 lines are prohibited. Extract them.
  • "I'll instruct the agent to parse the file line-by-line in text" -> WRONG. Use "Script-First".

2. Capabilities

  • Audit: Detect gaps (missing Red Flags, inline blocks > 12 lines, poor CSO, weak language) using analyze_gaps.py.
  • Execution Policy Audit: Detect missing Execution Mode, Script Contract, Safety Boundaries, and Validation Evidence sections.
  • Security Remediation: Fix vulnerabilities flagged by skill-validator (e.g., curl | bash, secrets, weak permissions).
  • Plan: Propose specific content improvements using references/refactoring_patterns.md.
  • Execute: Apply refactoring patterns to upgrade the skill.

2.5. Execution Mode

  • Mode: hybrid
  • Rationale: gap triage and refactoring decisions are prompt-driven, while gap detection is script-driven.

2.6. Script Contract

  • Primary Command: python3 scripts/analyze_gaps.py <target-skill-path> [--json]
  • Inputs: target skill path + optional output mode.
  • Outputs: structured gap list and pass/fail status.
  • Failure Semantics: non-zero exit when gaps exist (for deterministic gate behavior).

2.7. Safety Boundaries

  • Scope: apply edits only to explicitly selected target skill.
  • Default Exclusions: do not refactor unrelated skills or global docs by default.
  • Destructive Actions: full-file overwrite is prohibited unless explicitly requested and reviewed.

Read the full file on GitHub · 150 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 150 lines · 25 tokens per session scan B f2417392b5a8

Subscribe to this mod's changes

skill-enhancer is a skill published in the GitHub repository MatrixFounder/Agentic-development (5 stars, last pushed 19d ago), licensed Apache-2.0. It adds 25 tokens to every session and 2,292 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it B with 2 findings (downloads and executes remote code, makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

argent-tv-interact

Control and inspect TV apps via argent — Apple TV (tvOS), Android TV (leanback), and Amazon Fire TV (Vega). Boot the target, read focus, navigate with the D-pad remote, type, screenshot, and on Vega debug the JS runtime (evaluate, console logs, network inspector). Use when a task targets a TV (runtimeKind "tv", or…

software-mansion/argent · 107 tokens

review-offered-task

Review a task that has been offered to you and decide whether to accept or reject it.

desplega-ai/agent-swarm · 22 tokens

company-hiring-intelligence

Reverse-engineer what a company is building by scraping their job postings, careers page, LinkedIn Jobs, and engineering blog using TinyFish web agents. Use whenever a user wants to understand a company's strategic direction from hiring signals, do competitive intelligence, figure out a tech stack from job…

tinyfish-io/tinyfish-cookbook · 169 tokens

文档协作

引导用户通过结构化的文档共同编写工作流程。当用户想撰写文档、提案、技术规范、决策文档或类似结构化内容时使用。该工作流程帮助用户高效传递上下文,通过迭代优化内容,并验证文档对读者有效。当用户提到写文档、创建提案、起草规范或类似文档任务时触发。.

Tencent/WeKnora · 95 tokens

aidd-dev:08:for-sure

Iterative agent loop that tracks attempts and retries until a success condition is met. Use when the user says "for sure", "make sure", "keep trying until", "loop until done", "don't stop until", or needs guaranteed completion of a task with explicit success criteria.

ai-driven-dev/framework · 66 tokens

Swift Performance Optimization Skill

Use when investigating measured Swift or Apple-platform regressions in CPU, memory, launch, scrolling, animation hitches, image processing, energy, networking, or concurrency, or when designing performance tests and Instruments experiments. Do not use for speculative micro-optimization, ordinary refactoring, or a…

termio-sh/termio · 68 tokens