gsd-debugger
01Skill Claude CodeCodex
Investigates bugs using scientific method, manages debug sessions, handles checkpoints. Spawned by /gsd:debug orchestrator or diagnose-issues workflow.
58 tagged hypothesis testing, measured the same way as everything else here.
Browse within: product-discovery 31prd 16product-management 16annotations 12feedback-loop 12llmops 12regression-testing 12systematic-evaluation 12cockpit 6
Skill Claude CodeCodex
Investigates bugs using scientific method, manages debug sessions, handles checkpoints. Spawned by /gsd:debug orchestrator or diagnose-issues workflow.
Skill Claude CodeCodex
Colour, font and design-token rules for the docs site — when hex is banned, when it is correct, and which surfaces ignore the theme. Use when styling docs components or editing CSS under docs/src.
Skill Claude CodeCodex
Manage git worktrees with their own dev ports, symlinked .env files, playground, and simulations. Use when the user asks to create, set up, list, enter, or remove a worktree.
Skill Claude CodeCodex
Design, run, and analyze AI test suites on Rhesis — explore endpoints, build test foundations from a spec, create requirements and metrics, execute tests, and analyze results. Use when testing an AI endpoint, pasting a Product Requirements Document (PRD) or product spec, or working with Rhesis via MCP.
Skill Claude CodeCodex
Use this skill whenever a researcher wants to test, validate, stress-test, or falsify a research idea or hypothesis — especially in AI/ML/deep learning. Trigger on phrases like "I have an idea," "would this work," "test this hypothesis," "sanity check my idea," "what's wrong with this idea," "review my results," "is…
Skill Claude CodeCodex
Defines 4 Risks confidence thresholds, OST hierarchy levels, Knowledge Pyramid tiers, and state design requirements. Use when evaluating user stories, setting confidence scores, referencing OST levels, scoping MVP, or determining validation sufficiency.
Skill Claude CodeCodex
Orchestrate PRD creation from validated hypotheses — standard PRD output with 4 Risks confidence and hypothesis traceability.
Skill Claude CodeCodex
Orchestrate hypothesis validation through type-appropriate methods — prototypes, code analysis, market research, and expert review.
Skill Claude CodeCodex
Investigates bugs using scientific method, manages debug sessions, handles checkpoints. Spawned by /gsd:debug orchestrator or diagnose-issues workflow.
Skill Claude CodeCodex
Author: Dr. Aneesh Joseph Implementation: Claude (Anthropic) Version: 2.2 | December 2025.
Skill Claude CodeCodex
Use when debugging or changing behavior based on an uncertain causal hypothesis that the next action can test.
Skill Claude CodeCodex
Defines structural design artifact formats — information architecture, user flows, content model, brand direction, Visual Tokens, and AI interaction model. Use when creating or reviewing design artifacts that precede prototype generation.
Skill Claude CodeCodex
Manages hypothesis lifecycle, enforces validation criteria, time budgets, and confidence scoring rules. Use when creating hypotheses, updating confidence scores, setting validation criteria, handling timeouts, or recording validation results.
Skill Claude CodeCodex
Defines 4 Risks confidence thresholds, OST hierarchy levels, Knowledge Pyramid tiers, and state design requirements. Use when evaluating user stories, setting confidence scores, referencing OST levels, scoping MVP, or determining validation sufficiency.
Skill Claude CodeCodex
A literature wiki for a crux research vault — Andrej Karpathy's LLM-wiki pattern applied to a crux project. The PI curates immutable sources under raw/; you (the agent) compile them into a persistent, interlinked wiki/ of background, prior methods, SOTA, baselines, datasets, and definitions, and draw on it to ask…
Skill Claude CodeCodex
An agentic research companion — a scientific-method lab notebook for a research program. Organize work as a tree of Questions (what we don't know) and falsifiable Hypotheses (testable leaves), each with pre-registered verifiables and findings; a deterministic engine derives verdicts, rolls them up, and trips a human…
Skill Claude CodeCodex
Evolve crux itself — add a capability or fix a recurring flaw in the crux tool, end to end: ideate → build → validate → ship. Turns a feature idea or a "crux keeps doing X" annoyance into a signed-off PRD, a tests-first implementation, a hard validation gate (selftest green · stdlib-only · existing vaults still load ·…
Skill Claude CodeCodex
A collection of 30 tools for statistical analysis, including tests, regression models, and time-series methods. It is intended for examining data and checking whether observed patterns are meaningful.