Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/mckruz/claude-code-sdlc/eval-buildernpx skills add MCKRUZ/claude-code-sdlc --skill eval-buildergit clone --depth 1 https://github.com/MCKRUZ/claude-code-sdlcWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00078 | $0.00658 |
| Opus 5 | $0.00039 | $0.00329 |
| Sonnet 5 | $0.00016 | $0.00132 |
| Haiku 4.5 | $0.00008 | $0.00066 |
Grade A, and why
eval-builder scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 44 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Eval builder
For LLM-powered work, "tests pass" is not sufficient verification. The golden set is the spec's acceptance criteria; CI runs it like tests, and a regression in it blocks a merge like a failing test. This skill builds one that actually tracks quality.
Procedure (grounded in Anthropic's agent-eval guidance)
- Start from real failures, small. 20–50 tasks is a strong start — drawn from the manual pre-release checks and the bug/support queue, not invented in the abstract. Each task is one a second person would grade the same way (unambiguous, with a reference answer).
- Pick graders deterministic-first. Prefer a
state_check(did it reach the right end state?) ortranscript_constraint(e.g. finished in ≤ N turns) over an LLM judge. Use anllm_rubriconly for what rules can't capture (tone, explanation quality), and grade the output, not the path — don't assert an exact tool-call sequence; agents find valid routes you didn't predict. - Compose multidimensional success where needed (state + transcript + rubric), with partial credit for multi-part tasks instead of all-or-nothing.
- Write it under
eval-datasets/specs/<feature>/, referencing the spec file, versioned:eval-datasets/specs/<feature>/golden-set.yamlwithspec: specs/NNNN-<feature>.md(usekit/eval-datasets/golden-set.template.yamlas the shape). Set the threshold the spec requires (e.g. ">= 95%"). - Plan for variance. Agents are stochastic — the suite runs multiple trials per task from a clean, isolated environment (so prior-trial state can't leak or be gamed). Note the trial count.
Calibrate before you gate
- Validate the LLM judges against a few human-graded cases before trusting them; the regression trip-wire (~±3% is a practitioner starting point, not a constant) must be calibrated to the suite's measured variance before it becomes a required check.
Done when
- A versioned
golden-set.yamlexists undereval-datasets/specs/<feature>/(referencing the spec file) with 20–50 grounded cases. - Graders are deterministic where possible; LLM rubrics grade output, not trajectory.
- The threshold and trial count are set; the judge has been sanity-checked against human grades.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 44 lines · 78 tokens per session scan A 08683794e760
eval-builder is a skill published in the GitHub repository MCKRUZ/claude-code-sdlc (4 stars, last pushed 4d ago), licensed MIT. It adds 78 tokens to every session and 658 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
mcporter
List, auth, and call MCP servers/tools from the terminal.
agent-messaging
Send and receive cryptographically signed messages between AI agents using the Agent Messaging Protocol (AMP). Use when the user asks to "send a message to an agent", "check agent inbox", "message another agent", "reply to a message", "notify an agent", or any inter-agent communication task.
📝 任务完成后归档
重要提醒: 每次完成复杂调试或开发任务后,主动执行此流程! 将学到的经验归档为 skill,供以后参考。不要等用户提醒。.
oracle
Best practices for using the oracle CLI (prompt + file bundling, engines, sessions, and file attachment patterns).
agent-mode
Unified tool for managing agent LLM modes (add, remove, update, list, switch).
agento11y-prod-setup
Sets up production evaluation and guardrails for a DEPLOYED AI agent in Grafana Agent Observability, grounded in the agent's own code and its real ingested traffic. The judgment layer on top of the agento11y skill: it reads the agent's source (system prompt, tools, entrypoint) AND samples its live traffic via gcx…