artifact-formalizer

artifact-formalizer is a skill for Claude Code, Codex from MatrixFounder/Agentic-development. It costs 126 tokens per session (4,429 once invoked), scanned A, original, Apache-2.0.

A writing and review skill for technical specifications such as tasks, architecture documents, plans, and issue records.

In plain words
What is it for?
Use it before writing a specification to structure requirements clearly, or afterward to audit an existing document for unclear wording and unsupported judgments.
Why use it?
It helps prevent requirements from being buried in long, vague, or hard-to-check prose.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/matrixfounder/agentic-development/artifact-formalizer
Any agent
npx skills add MatrixFounder/Agentic-development --skill artifact-formalizer
Clone the repo
git clone --depth 1 https://github.com/MatrixFounder/Agentic-development

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for artifact-formalizer

README.md
[![agentmods](https://agentmods.dev/badge/skills/matrixfounder/agentic-development/artifact-formalizer.svg)](https://agentmods.dev/skills/matrixfounder/agentic-development/artifact-formalizer)
Your own site
<a href="https://agentmods.dev/skills/matrixfounder/agentic-development/artifact-formalizer"><img src="https://agentmods.dev/badge/skills/matrixfounder/agentic-development/artifact-formalizer.svg" alt="Measured on agentmods" height="20"></a>
Per session 126 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 4,429 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00126 $0.04429
Opus 5 $0.00063 $0.02214
Sonnet 5 $0.00025 $0.00886
Haiku 4.5 $0.00013 $0.00443

Measured today against content hash 721456843f49, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

artifact-formalizer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agent/skills/artifact-formalizer/SKILL.md · 286 lines

How it starts

The opening of the file, as written. The whole thing — 286 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Artifact formalizer (specification register)

Purpose

A reader of an essay-register artifact performs two passes: one to locate the requirement, one to decide whether a given sentence carried one. This skill removes the second pass. It has two modes.

Mode When Instrument Owns
A — Authoring before and during writing references/authoring-contract.md the defect is not written
B — Audit on an existing document scripts/scan_register.py + a reading pass what Mode A missed

Mode A prevents the defect; Mode B measures what Mode A missed.

Why in that order. Defective prose measured 5.1% of one corpus's words (731 of 14,288), while removing it required reading and editing all 14,288. Register also varied by authoring model on the same repository, so an unwritten standard is not a standard. references/measurement-baseline.md §5 carries both figures.

Scope boundary — what this is NOT

This is not a jargon or terminology tool. Domain terms are precise and stay verbatim: singleflight, RTM, дедлайн, throttle. Translating engineering vocabulary into business language for customer-facing documents is a different problem with a different audience, and this skill does not attempt it.

1. Red Flags (Anti-Rationalization)

A definition list, not a table: the reality column is prose, and documentation-standards §5.1 prescribes converting such a column rather than widening it.

  • "I'll write it, then run the formalizer." That is the expensive order. §Purpose gives the measured cost. Mode A first.
  • "The scan is clean, so the text is clean." Every zero is reported next to what the detector actually saw. Read DIAGNOSTICS: a corpus whose longest sentence equals the limit was written for the gate.
  • "I'll add a pattern for this new phrasing." First ask whether a test in the authoring contract already forbade it. If it did, the lexicon gains a faster detector and nothing else changes. If it did not, the contract is amended (§6).
  • "«Шов», «нога», «мост» — это термины проекта." A term appears in ARCHITECTURE.md, a public API, or a cited standard. Run --terms docs/ARCHITECTURE.md and let the scanner apply that test.
  • "I formalized the section that was quoted at me." The pass that produced this skill's own worked example did exactly that. It left seventeen occurrences of one metaphor in the same task set. Coverage is per section: --sections.
  • "The whole document reads badly — I'll rewrite it." Conforming sentences stay verbatim. Over-rewriting introduces errors that were not there.
  • "I'll shorten it by trimming the requirement." Register changes, substance does not. Numbers, identifiers, obligations and scope limits survive verbatim.
  • "This word is fine — I'll widen the threshold." A failing scan is fixed in the prose. Thresholds move only when a measurement moves them.
  • "It flagged a false positive, so the rule is wrong." The scanner is advisory by design. Judge the finding; info exists because the author decides.

Read the full file on GitHub · 286 lines

Files

What ships with it

60 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. today Changed · +4 tokens per session 721456843f49
  2. 4d ago First seen · 286 lines · 122 tokens per session scan A bc0e788b4446

Subscribe to this mod's changes

artifact-formalizer is a skill published in the GitHub repository MatrixFounder/Agentic-development (5 stars, last pushed yesterday), licensed Apache-2.0. It adds 126 tokens to every session and 4,429 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

Learn what defines effective BDD scenarios

Interactive guidance on writing complete, effective BDD scenarios for story-flow.

Intai/story-flow · 22 tokens

add-tests

Generates tests for existing code. Analyzes the target function, method, or class to identify the happy path, error cases, and edge cases, then writes test cases following the project's testing framework and naming conventions. Invoked when the user asks to add tests, write tests, cover code, or increase coverage.

soulcodex/agentic · 67 tokens

tdd-implementation

Use when implementing any code change for a task – new behavior, a bugfix, or a revision after review – before writing the production code. Defines the failing-test-first cycle, public-seam testing, vertical slices, scaffold and infra handling, and counters to common excuses for tests-after-code.

bartoszarendt/agenticloop · 64 tokens

test-writer

Writes tests that fail before a fix and pass after it.

rysh-ai/rysh-cli-code · 16 tokens

test-generation

Generate and run unit/integration tests TDD-style across Python, Node, and .NET. Use when adding tests to untested code, implementing a feature test-first, or finding coverage gaps.

jnotsknab/mux-swarm · 42 tokens

argent-tv-interact

Control and inspect TV apps via argent — Apple TV (tvOS), Android TV (leanback), and Amazon Fire TV (Vega). Boot the target, read focus, navigate with the D-pad remote, type, screenshot, and on Vega debug the JS runtime (evaluate, console logs, network inspector). Use when a task targets a TV (runtimeKind "tv", or…

software-mansion/argent · 107 tokens