hypothesis testing skills

58 tagged hypothesis testing, measured the same way as everything else here.

Browse within: product-discovery 31prd 16product-management 16annotations 12feedback-loop 12llmops 12regression-testing 12systematic-evaluation 12cockpit 6

gsd-debugger

01

allgpt-co/QuickVoice

Skill Claude CodeCodex

Investigates bugs using scientific method, manages debug sessions, handles checkpoints. Spawned by /gsd:debug orchestrator or diagnose-issues workflow.

492 20d ago A 35 tokens original MIT

docs-design-tokens

02

rhesis-ai/rhesis

Skill Claude CodeCodex

Colour, font and design-token rules for the docs site — when hex is banned, when it is correct, and which surfaces ignore the theme. Use when styling docs components or editing CSS under docs/src.

391 4d ago A 46 tokens

worktree

03

rhesis-ai/rhesis

Skill Claude CodeCodex

Manage git worktrees with their own dev ports, symlinked .env files, playground, and simulations. Use when the user asks to create, set up, list, enter, or remove a worktree.

391 4d ago A 46 tokens

rhesis

04

rhesis-ai/rhesis

Skill Claude CodeCodex

Design, run, and analyze AI test suites on Rhesis — explore endpoints, build test foundations from a spec, create requirements and metrics, execute tests, and analyze results. Use when testing an AI endpoint, pasting a Product Requirements Document (PRD) or product spec, or working with Rhesis via MCP.

391 4d ago A 67 tokens

saidwivedi/research-skills

Skill Claude CodeCodex

Use this skill whenever a researcher wants to test, validate, stress-test, or falsify a research idea or hypothesis — especially in AI/ML/deep learning. Trigger on phrases like "I have an idea," "would this work," "test this hypothesis," "sanity check my idea," "what's wrong with this idea," "review my results," "is…

40 2mo ago A 108 tokens original MIT

product-principles

06

shinpr/claude-code-discover

Skill Claude CodeCodex

Defines 4 Risks confidence thresholds, OST hierarchy levels, Knowledge Pyramid tiers, and state design requirements. Use when evaluating user stories, setting confidence scores, referencing OST levels, scoping MVP, or determining validation sufficiency.

10 3d ago A 49 tokens original MIT

recipe-define

07

shinpr/claude-code-discover

Skill Claude CodeCodex

Orchestrate PRD creation from validated hypotheses — standard PRD output with 4 Risks confidence and hypothesis traceability.

10 3d ago A 27 tokens original MIT

recipe-validate

08

shinpr/claude-code-discover

Skill Claude CodeCodex

Orchestrate hypothesis validation through type-appropriate methods — prototypes, code analysis, market research, and expert review.

10 3d ago A 26 tokens original MIT

gsd-debugger

09

ctsstc/get-shit-done-skills

Skill Claude CodeCodex

Investigates bugs using scientific method, manages debug sessions, handles checkpoints. Spawned by /gsd:debug orchestrator or diagnose-issues workflow.

9 7mo ago A 35 tokens

Claude

10

Dr-AneeshJoseph/Prism

Skill Claude CodeCodex

Author: Dr. Aneesh Joseph Implementation: Claude (Anthropic) Version: 2.2 | December 2025.

7 8mo ago A 0 tokens original MIT

bet-on-it

11

Lum1104/bet-on-it

Skill Claude CodeCodex

Use when debugging or changing behavior based on an uncertain causal hypothesis that the next action can test.

5 1mo ago A 23 tokens original MIT

blueprint-standards

12

shinpr/nautilus

Skill Claude CodeCodex

Defines structural design artifact formats — information architecture, user flows, content model, brand direction, Visual Tokens, and AI interaction model. Use when creating or reviewing design artifacts that precede prototype generation.

4 3d ago A 45 tokens original MIT

shinpr/nautilus

Skill Claude CodeCodex

Manages hypothesis lifecycle, enforces validation criteria, time budgets, and confidence scoring rules. Use when creating hypotheses, updating confidence scores, setting validation criteria, handling timeouts, or recording validation results.

4 3d ago A 45 tokens original MIT

product-principles

14

shinpr/nautilus

Skill Claude CodeCodex

Defines 4 Risks confidence thresholds, OST hierarchy levels, Knowledge Pyramid tiers, and state design requirements. Use when evaluating user stories, setting confidence scores, referencing OST levels, scoping MVP, or determining validation sufficiency.

4 3d ago A 49 tokens original MIT

crux-wiki

15

mehdiforoozandeh/crux

Skill Claude CodeCodex

A literature wiki for a crux research vault — Andrej Karpathy's LLM-wiki pattern applied to a crux project. The PI curates immutable sources under raw/; you (the agent) compile them into a persistent, interlinked wiki/ of background, prior methods, SOTA, baselines, datasets, and definitions, and draw on it to ask…

4 2d ago A 219 tokens original MIT

crux

16

mehdiforoozandeh/crux

Skill Claude CodeCodex

An agentic research companion — a scientific-method lab notebook for a research program. Organize work as a tree of Questions (what we don't know) and falsifiable Hypotheses (testable leaves), each with pre-registered verifiables and findings; a deterministic engine derives verdicts, rolls them up, and trips a human…

4 2d ago A 166 tokens original MIT

evolve-crux

17

mehdiforoozandeh/crux

Skill Claude CodeCodex

Evolve crux itself — add a capability or fix a recurring flaw in the crux tool, end to end: ideate → build → validate → ship. Turns a feature idea or a "crux keeps doing X" annoyance into a signed-off PRD, a tests-first implementation, a hard validation gate (selftest green · stdlib-only · existing vaults still load ·…

4 2d ago A 218 tokens original MIT

mcp-stats-engine-zx

18

wzx11223344/mcp-stats-engine

Skill Claude CodeCodex

A collection of 30 tools for statistical analysis, including tests, regression models, and time-series methods. It is intended for examining data and checking whether observed patterns are meaningful.

0 1mo ago A 7 tokens original MIT