Skill Claude CodeCodex
Determines whether validation has reached saturation — no new weaknesses or failure modes being discovered. Used by all 5 campaigns as termination signal.
Research Artifact Stress-Testing Engine — five-campaign adversarial validation producing weakness-annotated verification reports
Skill Claude CodeCodex
Determines whether validation has reached saturation — no new weaknesses or failure modes being discovered. Used by all 5 campaigns as termination signal.
Skill Claude CodeCodex
Synthesize breakpoints across dimensions into a coherent validity envelope for a claim.
Skill Claude CodeCodex
Map the complete validity envelope of a claim across all relevant dimensions, synthesizing breakpoints into a bounded region.
Skill Claude CodeCodex
Deep web full-text retrieval via Apify RAG browser. Import of web-browsing/web-research skill. Full page content for substantive analysis.
Skill Claude CodeCodex
Quick web scanning for landscape understanding. Import of web-browsing/web-search skill. Snippets only — no conclusions from snippets alone.
Skill Claude CodeCodex
Strategy: Pearl Three-Step counterfactual — Abduction (fit model to evidence), Action (intervene on factor), Prediction (derive counterfactual outcome).
Skill Claude CodeCodex
Tactic: Full attack lifecycle — threat surface enumeration, attack vector generation, systematic probing, and finding aggregation across all surfaces.
Skill Claude CodeCodex
Evaluate the probability of sufficiency (PS) for a causal factor — would this factor alone be enough to produce the conclusion?
Skill Claude CodeCodex
Tactic: List all factors, remove one at a time, assess conclusion stability, rank factors by load-bearing importance.
Skill Claude CodeCodex
Strategy: AI-safety systematic probing — enumerate all threat surfaces, generate attack vectors per surface, execute probes, and aggregate findings across the full attack space.
Skill Claude CodeCodex
Strategy: Williamson-style precise thought experiments — construct carefully specified counterfactual scenarios to test whether conclusions depend on contingent features.
Skill Claude CodeCodex
Enumerate all attackable surfaces of an artifact — logical, empirical, methodological, social, and practical dimensions.
Skill Claude CodeCodex
Synthesizes findings from a completed campaign into typed verdict reports. Produces DebateVerdict, RedTeamReport, FailureAnticipationReport, CounterfactualMap, or AdversarialStressReport depending on campaign. Also supports cross-campaign StressTestSummary.
Skill Claude CodeCodex
Classifies discovered weaknesses into severity tiers (fatal/major/minor/cosmetic) with structured justification and exploitability assessment.
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: