strat-creator: Skill for Claude Code

.claude/skills/strategy-testability-review/SKILL.md

strategy-testability-review is a skill for Claude Code from opendatahub-io/strat-creator. It costs 27 tokens per session (879 once invoked), scanned A, original, Apache-2.0.

A review of strategy documents that checks whether proposed features can be proven to work through concrete tests.

In plain words
What is it for?
Use it to assess acceptance criteria, measurable outcomes, edge cases, and prior reviews in strategy files and their original requirement documents.
Why use it?
It finds vague success conditions and missing edge cases before development, when they are easier to clarify.

Skill for Claude Code

Written for Claude Code: allowed-tools in frontmatter. Also seen: model in frontmatter.

This is opendatahub-io/strat-creator's own configuration. It tells Claude Code how to work on strat-creator itself, so it is not a mod to install elsewhere. Copy it as a starting point and replace the rules that are about this project. Everything strat-creator configures →

Reuse

Borrowing it

Nothing to install: this file belongs to opendatahub-io/strat-creator. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.

Copy the file
curl -O https://raw.githubusercontent.com/opendatahub-io/strat-creator/main/.claude/skills/strategy-testability-review/SKILL.md
Clone the repo
git clone --depth 1 https://github.com/opendatahub-io/strat-creator

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for strategy-testability-review

README.md
[![agentmods](https://agentmods.dev/badge/skills/opendatahub-io/strat-creator/strategy-testability-review/github.svg)](https://agentmods.dev/skills/opendatahub-io/strat-creator/strategy-testability-review)
Your own site
<a href="https://agentmods.dev/skills/opendatahub-io/strat-creator/strategy-testability-review"><img src="https://agentmods.dev/badge/skills/opendatahub-io/strat-creator/strategy-testability-review/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for strategy-testability-review

Your own site · 80×15
<a href="https://agentmods.dev/skills/opendatahub-io/strat-creator/strategy-testability-review"><img src="https://agentmods.dev/badge/skills/opendatahub-io/strat-creator/strategy-testability-review.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 27 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 879 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00027 $0.00879
Opus 5 $0.00014 $0.00439
Sonnet 5 $0.00005 $0.00176
Haiku 4.5 $0.00003 $0.00088

Measured 11d ago against content hash ac55409754be, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-11, from the pricing page.

Security

Grade A, and why

strategy-testability-review scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/skills/strategy-testability-review/SKILL.md · 60 lines

How it starts

The opening of the file, as written. The whole thing — 60 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are a test engineer reviewing refined strategy features. Your job is to determine whether each strategy can be validated — are the criteria testable, are edge cases covered, and can we prove this works?

Inputs

Check if strategy files exist in local/strat-tasks/. If they do, use local mode:

  • Read strategies from local/strat-tasks/
  • Read RFE originals from local/strat-originals/
  • Read prior reviews from local/strat-reviews/

Otherwise use CI mode:

  • Read strategies from artifacts/strat-tasks/
  • Read RFE originals from artifacts/rfe-tasks/
  • Read prior reviews from artifacts/strat-reviews/

If $ARGUMENTS contains a strategy key (e.g., RHAISTRAT-133), review only that strategy. Otherwise review all strategies in the directory.

Cross-reference against the source RFEs for the original acceptance criteria. If this is a re-review (prior review files exist), read them.

What to Assess

For each strategy:

  1. Are acceptance criteria testable? Can each criterion be verified with a concrete test? "Users can do X" is testable. "System is reliable" is not.
  2. Are success criteria measurable? If the RFE says ">80% reduction in tokens," can we measure that? What's the baseline?
  3. What edge cases are missing? Failure modes, boundary conditions, concurrent access, large-scale scenarios, backwards compatibility with existing deployments.
  4. What's the test strategy? Unit tests, integration tests, e2e tests — what's needed to validate this? Are there components that are hard to test (external dependencies, multi-cluster scenarios)?
  5. Are non-functional requirements testable? Performance benchmarks, scalability limits, security requirements — can we write tests for these?
  6. Are acceptance criteria structured? Given/When/Then format with "measured by" clauses is a positive signal — it forces testable specificity. Bullet-list criteria without verification methods are a gap. Flag criteria that use vague language: "works correctly", "users can easily...", or capability lists without verification.
  7. Are NFRs quantified with numeric thresholds? Each non-functional requirement should have a measurable target (latency in ms, throughput in requests/sec, error rate percentage). Generic NFRs like "good performance", "secure access", or "high availability" are not testable. Missing NFRs entirely for L/XL strategies is a testability gap.
  8. Are NFR metrics grounded in a cited source? Every numeric threshold in the NFR section must cite where it came from — the RFE, a specific architecture context doc, or Staff Engineer / SME Input. If an NFR states a number (e.g., "< 500ms response time", "1000 RPS", "500 concurrent users") without citing a source, flag it as an ungrounded metric. Architectural facts like replica counts, TLS versions, and HPA ranges from the platform docs are valid. Invented performance or scalability targets with no source are not — they should be open questions, not stated as requirements.

Read the full file on GitHub · 60 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 11d ago First seen · 60 lines · 27 tokens per session scan A ac55409754be

Subscribe to this mod's changes

strategy-testability-review is a skill published in the GitHub repository opendatahub-io/strat-creator (2 stars, last pushed yesterday), licensed Apache-2.0. It adds 27 tokens to every session and 879 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

research-engineer

An uncompromising Academic Research Engineer. Operates with absolute scientific rigor, objective criticism, and zero flair. Focuses on theoretical correctness, formal verification, and optimal implementation across any required technology.

davila7/claude-code-templates · 43 tokens

tika-eval-compare

Compare extracts from two Tika builds over a corpus to detect regressions in content, encoding, exceptions, and embedded-document handling. Use for "compare before/after extracts", "eval this change against the corpus".

apache/tika · 50 tokens

neuron-evaluation-engineer

Create and run AI evaluations with datasets, assertions, and output drivers in Neuron AI. Use this skill whenever the user mentions evaluation, testing AI systems, creating evaluators, dataset-driven testing, assertion-based validation, or wants to measure AI system performance. Also trigger for tasks involving…

neuron-core/neuron-ai · 77 tokens

jetson-validate-image

Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.

NVIDIA/skills · 50 tokens

atmos-validation

Validate Atmos projects, components, arbitrary JSON Schema inputs, EditorConfig, and GitHub Actions; use affected-file selection and native CI annotations.

cloudposse/atmos · 31 tokens

skill-benchmark

Benchmark AI skill effectiveness by measuring implementation quality against legacy constraints.

HoangNguyen0403/agent-skills-standard · 16 tokens