chronicle-mcp: Skill for Claude Code

.agents/skills/skill-tdd/SKILL.md

skill-tdd is a skill for Claude Code, Codex from loerei/chronicle-mcp. It costs 17 tokens per session (2,565 once invoked), scanned A, original, MIT.

A test-driven development guide for testing and evaluating coding-agent skills and policy files. TDD means checking behavior with tests before accepting an implementation, including both cases that should fail and cases that should pass.

In plain words
What is it for?
It helps create baseline and treated runs, test multiple domains and failure patterns, compare results, and decide whether a skill is ready for broader use.
Why use it?
It exposes missing rules, incorrect decisions, and overly strict behavior through repeatable test cycles.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one. Also seen: mentions subagents; positional $N argument; installed under .agents/ (shared by several agents).

This is loerei/chronicle-mcp's own configuration. It tells Claude Code and Codex how to work on chronicle-mcp itself, so it is not a mod to install elsewhere. Copy it as a starting point and replace the rules that are about this project. Everything chronicle-mcp configures →

Reuse

Borrowing it

Nothing to install: this file belongs to loerei/chronicle-mcp. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.

Copy the file
curl -O https://raw.githubusercontent.com/loerei/chronicle-mcp/main/.agents/skills/skill-tdd/SKILL.md
Clone the repo
git clone --depth 1 https://github.com/loerei/chronicle-mcp

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for skill-tdd

README.md
[![agentmods](https://agentmods.dev/badge/skills/loerei/chronicle-mcp/skill-tdd/github.svg)](https://agentmods.dev/skills/loerei/chronicle-mcp/skill-tdd)
Your own site
<a href="https://agentmods.dev/skills/loerei/chronicle-mcp/skill-tdd"><img src="https://agentmods.dev/badge/skills/loerei/chronicle-mcp/skill-tdd/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for skill-tdd

Your own site · 80×15
<a href="https://agentmods.dev/skills/loerei/chronicle-mcp/skill-tdd"><img src="https://agentmods.dev/badge/skills/loerei/chronicle-mcp/skill-tdd.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 17 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,565 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00017 $0.02565
Opus 5 $0.00009 $0.01282
Sonnet 5 $0.00003 $0.00513
Haiku 4.5 $0.00002 $0.00257

Measured 9d ago against content hash d9fd4b70d856, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-09, from the pricing page.

Security

Grade A, and why

skill-tdd scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agents/skills/skill-tdd/SKILL.md · 125 lines

How it starts

The opening of the file, as written. The whole thing — 125 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Skill TDD

Test and benchmark skills (SKILL.md), subdocs, and policies (AGENTS.md) using red-green-regression cycles, 4-axis coverage, single-call parallel dispatch, and hypothesis ledgers.

Directives

  1. Red Baseline First: MUST run a baseline subagent without the skill on negative fixtures to verify failure (Red) before testing candidate versions.
  2. Double-Sided Fixtures: MUST test both a Negative Fixture (agent must block/refactor) and a Positive Fixture (agent must clear/approve) to prevent false-positive hyper-critique.
  3. 4-Axis Coverage: Test suites MUST cover Directives, Decision Branches, Anti-Patterns, and Domain Diversity (minimum 2–3 distinct domains) per REFERENCE.md.
  4. Test Matrix Budget:
    • Fast Dev Loop: $N_{\text{tests}} \ge 2$, $N_{\text{runs}} \in [1, 3]$ per test.
    • Graduation Gate: $N_{\text{tests}} \ge 3\text{--}4$ diverse domain fixtures, $N_{\text{runs}} \ge 5\text{--}10$ with fresh 0-prior-context subagents.
    • Graduation Criteria: $\Delta_{\text{Suite}} \ge +60%$, 100% positive clearance, and 0 cheat flags.
  5. Single-Call Parallel Matrix Dispatch: MUST launch all independent test trials (baseline and treated runs across all fixtures) in a single concurrent batch via invoke_subagent(Subagents=[...]). NEVER dispatch trials sequentially across multiple turns. Tag each subagent clearly: Role: "[Baseline/Treated] | <Fixture> | Run <K>".
  6. Spec Decoupling & Anti-Snooping Audit:
    • TEST_SPEC.md and HYPOTHESIS.md MUST reside in .scratch/<skill>-versions/specs/ and NEVER inside fixture code directories (.scratch/fixtures/<name>/).
    • Evaluators MUST inspect subagent tool calls in transcripts. If a baseline subagent reads repository SKILL.md files or ANY subagent reads TEST_SPEC.md/HYPOTHESIS.md, abort immediately with:
      Overfitting cheat detected: Subagent snooped <path>.
  7. Immutable Prompt Contract: MUST copy prompts verbatim from TEST_SPEC.md. NEVER inject hints, line numbers, or leading context.
  8. Zero Prior Context: Spawn every test run using a fresh subagent with 0 prior context. NEVER reuse subagent conversation threads.
  9. Sandbox Delta Versioning:
    • NEVER edit production skills/policies directly on test failure.
    • Create a sandbox (.scratch/<skill>-versions/) with candidate files (SKILL.v1.md) and an append-only HYPOTHESIS.md (NEVER delete entries).
    • Point test subagents to candidate versions. Graduate to production only after meeting the $\Delta_{\text{Suite}}$ threshold.
  10. Anti-Overfitting Cheats:
    • SKILL.md and REFERENCE.md MUST express domain-agnostic rules. NEVER hardcode test fixture entities, class names, or mock values.
    • Evaluators detecting leaked words or prompt hints MUST abort immediately with:
      Overfitting cheat detected: <reason>.
  11. Zero-Delta Anomaly: If 100% of runs yield Double-GREEN ($\Delta = 0$), classify as a Zero-Delta Failure (harden test fixture or prune redundant skill rules).
  12. Seam-Only Assertions: Verify behavior at public seams (verdict, scorecard rows, tool call constraints, stop boundary). NEVER assert on private reasoning traces.

Read the full file on GitHub · 125 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 9d ago First seen · 125 lines · 17 tokens per session scan A d9fd4b70d856

Subscribe to this mod's changes

skill-tdd is a skill published in the GitHub repository loerei/chronicle-mcp (0 stars, last pushed 4d ago), licensed MIT. It adds 17 tokens to every session and 2,565 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.