toolsmith: Skill for Claude Code

.claude/skills/next-milestone/SKILL.md

next-milestone is a skill for Claude Code from kayne-lee/toolsmith. It costs 47 tokens per session (808 once invoked), scanned A, original, MIT.

A repository workflow for continuing work from the previous coding session and advancing the next planned milestone, a defined stage of project work.

In plain words
What is it for?
Use it when asked to continue, start the next milestone, keep going, or report the project's current state.
Why use it?
It gives an agent a repeatable way to recover context, inspect unfinished changes, follow the existing plan, and verify completion.

Skill for Claude Code

Written for Claude Code: installed under .claude/.

This is kayne-lee/toolsmith's own configuration. It tells Claude Code how to work on toolsmith itself, so it is not a mod to install elsewhere. Copy it as a starting point and replace the rules that are about this project. Everything toolsmith configures →

Reuse

Borrowing it

Nothing to install: this file belongs to kayne-lee/toolsmith. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.

Copy the file
curl -O https://raw.githubusercontent.com/kayne-lee/toolsmith/main/.claude/skills/next-milestone/SKILL.md
Clone the repo
git clone --depth 1 https://github.com/kayne-lee/toolsmith

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for next-milestone

README.md
[![agentmods](https://agentmods.dev/badge/skills/kayne-lee/toolsmith/next-milestone.svg)](https://agentmods.dev/skills/kayne-lee/toolsmith/next-milestone)
Your own site
<a href="https://agentmods.dev/skills/kayne-lee/toolsmith/next-milestone"><img src="https://agentmods.dev/badge/skills/kayne-lee/toolsmith/next-milestone.svg" alt="Measured on agentmods" height="20"></a>
Per session 47 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 808 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00047 $0.00808
Opus 5 $0.00023 $0.00404
Sonnet 5 $0.00009 $0.00162
Haiku 4.5 $0.00005 $0.00081

Measured 7d ago against content hash 075cd8210038, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-07, from the pricing page.

Security

Grade A, and why

next-milestone scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 7d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/skills/next-milestone/SKILL.md · 62 lines

How it starts

The opening of the file, as written. The whole thing — 62 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Next milestone

Sessions on this repository start with no memory of the previous one. PROGRESS.md is the handoff.

Procedure

  1. Read PROGRESS.md bottom-up. The last entry is where work stopped. Read at least the last two — the second-to-last often explains a decision the last one depends on. Benchmark numbers live here; find the most recent one and the split it came from before changing anything.
  2. Read PLAN.md for the milestone about to be worked. The plan is fixed; if its scope turns out to be wrong, say so and propose an amendment rather than silently doing something else.
  3. Confirm the working tree is clean (git status). Uncommitted changes mean a previous session ended mid-work — read the diff before doing anything.
  4. Do the work. Tests are written alongside the code, not deferred. A milestone is not done because the code exists; it is done when its "Done when" criterion in PLAN.md is demonstrably met.
  5. Run the full check before committing:
    uv run ruff check . && uv run ruff format --check . && uv run mypy && uv run pytest
    
    All four must pass. If something fails and cannot be fixed within scope, stop and record the failure in PROGRESS.md rather than committing over it.
  6. Commit. Small, focused commits with imperative subjects (add schema tightening proposer, not Added schema tightening proposer). Never add a Co-Authored-By trailer or any AI-attribution footer.
  7. Append to PROGRESS.md: what was built, what was decided and why, the benchmark numbers with their split, what is blocked, and what the next session should start on.
  8. Close the milestone's issue with a comment summarizing the outcome.

Rules specific to this repository

  • Never touch the held-out split except to publish. Split.HELD_OUT requires publishing=True for exactly this reason. If you find yourself wanting to peek at held-out performance mid-iteration, that is the failure this project exists to avoid — don't.
  • Every accepted optimization has a recorded before/after. A rewritten description that was not benchmarked is not an improvement, it is a change. Record rejections too, with the number that caused them; the rejections are half the finding.
  • Report regressions in the first sentence. If an optimization strategy did not work, docs/results.md says so plainly. "This approach did not improve selection accuracy" is a legitimate result and a more credible repository than one where everything worked.
  • Do not generate or register executable tool code. Descriptions, schemas, and composition of existing implementations only. If a milestone seems to call for code generation, it does not — re-read PLAN.md.
  • Watch the tool-count tradeoff on macro synthesis. Each macro removes round trips and adds a candidate to choose among. Past some count, selection accuracy falls. Find that point and document it rather than assuming more is better.
  • API calls use claude-opus-5 with adaptive thinking. Cache the stable instruction prefix — the optimizer re-runs the same prompt shape many times per round and this is the dominant cost.

Read the full file on GitHub · 62 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 7d ago First seen · 62 lines · 47 tokens per session scan A 075cd8210038

Subscribe to this mod's changes

next-milestone is a skill published in the GitHub repository kayne-lee/toolsmith (0 stars, last pushed 1mo ago), licensed MIT. It adds 47 tokens to every session and 808 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.