Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add fledgeling-co/fledgeling-plugins --skill stocktakegit clone --depth 1 https://github.com/fledgeling-co/fledgeling-pluginsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/fledgeling-co/fledgeling-plugins/stocktake)<a href="https://agentmods.dev/skills/fledgeling-co/fledgeling-plugins/stocktake"><img src="https://agentmods.dev/badge/skills/fledgeling-co/fledgeling-plugins/stocktake.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00278 | $0.04312 |
| Opus 5 | $0.00139 | $0.02156 |
| Sonnet 5 | $0.00056 | $0.00862 |
| Haiku 4.5 | $0.00028 | $0.00431 |
Grade A, and why
stocktake scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 334 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Stocktake
A stocktake counts what is on the shelf against what the books say. Here the board is the books and the codebase is the shelf, and the count is the thing that finds the books wrong.
Three failure modes shape the whole design, and each of them produces a board that reads as healthy:
The card says done and nothing produces the data. A surface renders, the schema
validates, the suite is green, and no code ever wrote the value. spec-validation:spec-validation
exists because roughly half of a 110-ticket corpus shipped not-as-specified while
reading as complete.
The tests behind the card cannot fail. An assertion comparing a value with itself, a case satisfied by the wrong exception type, a fixture in a shape the product never stores. All three were found in a single session's own work, by an out-of-family reader, after the author had declared each one red-armed.
The evidence is authored by the party being judged. The worker writes the tests, the completion record and the comment that says it is finished. METR documents frontier agents editing tests and monkey-patching evaluators. So the ticket is the oracle and the worker's record is the defendant, and the order they are read in is load-bearing rather than stylistic.
Running as a Gemini model? Read gemini.md in this directory first, then follow this file with the overrides it names. It turns the sweep's categorical scopes into a filled quota ledger, makes every verdict row carry the lane's argv and its output, reads this skill's bounds — one judge, three cards per brief, the 50KB packet — back off the produced run, and converts spec-validation:spec-validation and clarify:clarify into phases that emit trace.md and decisions.md. Other models skip it.
The order that makes this work
Build the numbered requirement list from the description, every comment and every attached image BEFORE opening the completion record, the plan, or the diff.
Read in the other order and the diff tells you what to look for, so you find it. This
is the single rule most worth keeping; references/the-oracle-order.md carries why,
and the mechanics of reading image attachments as requirements rather than decoration.
What ships with it
13 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- gemini.md 19 KB
- references/column-policy.md 7.2 KB
- references/evidence.md 4.1 KB
- references/running-long.md 4.2 KB
- references/testing-adequacy.md 4.0 KB
- references/the-oracle-order.md 3.1 KB
- references/verification-lanes.md 5.0 KB
- scripts/board_ledger.py 10.0 KB runs code
- scripts/check_verified_gate.py 6.5 KB runs code
- scripts/gates.py 28 KB runs code
- scripts/locate_work.sh 2.3 KB runs code
- scripts/verify_queue.sh 2.6 KB runs code
- scripts/warrant_column.py 11 KB runs code
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- today Changed 14634fd5cee4
- 4d ago First seen · 334 lines · 278 tokens per session scan A e14624d3a527
stocktake is a skill published in the GitHub repository fledgeling-co/fledgeling-plugins (2 stars, last pushed yesterday), licensed MIT. It adds 278 tokens to every session and 4,312 once invoked, about $0.0014 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-04.
Other skills, from other repositories
scaffold-dotnet-test-project
MUST USE when an existing .NET test project was excluded from a .slnf/CI solution filter, disappeared from .sln/.slnx discovery, or lost its production ProjectReference; also for requests to set up, create, reuse, add, register, include, or repair a test project. Handles "tests pass directly but CI discovers zero"…
cy-execute-task
Implement and verify an existing CompozyOS spec task, then update its tracking. Excludes review remediation.
team-qa
Orchestrate the QA team through a full testing cycle. Coordinates qa-lead (strategy + test plan) and qa-tester (test case writing + bug reporting) to produce a complete QA package for a sprint or feature. Covers: test plan generation, test case writing, smoke check gate, manual QA execution, and sign-off report.
gh-issue
Size-audit, write, and split BanyanDB issues that somebody else or an automated TDD workflow can implement. Use whenever the user asks to file or revise an issue, decide whether an issue is too large, make an issue TDD-ready, turn a design into tickets, or split an umbrella into executable leaves. Do not draft or file…
implement-feature
Implement an approved feature plan with fresh-context slices, TDD, evidence, and PR-ready output.
conductor-implement
Execute tasks from a track's implementation plan following TDD workflow.