yogsoth-ai/stress-test

Research Artifact Stress-Testing Engine — five-campaign adversarial validation producing weakness-annotated verification reports

2Stars on the repository
110Mods indexed here, across every type
2mo agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

yogsoth-ai/stress-test

Skill Claude CodeCodex

Evaluate whether a derivation chain has reached a genuine contradiction, absurdity, or inconclusive state.

not rated 2 2mo ago A 26 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Systematically generate counterexamples (monsters) to a given claim using diverse heuristic strategies.

not rated 2 2mo ago A 22 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Generate counterexamples (monsters), attempt monster-barring, incorporate surviving counterexamples as lemma refinements (Lakatos method).

not rated 2 2mo ago A 31 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Campaign: Counterfactual reasoning to identify load-bearing factors. Core question: If key factors were different, would the conclusion still hold? Methods: Pearl SCM Three-Step, Lewis Possible Worlds, Tetlock & Belkin, PNS/PS.

not rated 2 2mo ago A 55 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Construct precise, internally consistent counterfactual scenarios where specified factors are altered, then reason about the resulting conclusion.

not rated 2 2mo ago A 30 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Legal adversarial structure — prosecution presents case, defense responds, evidence is cross-examined, judge delivers verdict. Emphasizes evidence quality and procedural rigor.

not rated 2 2mo ago A 39 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Classic triangular debate — Critic attacks, Defender responds, Judge adjudicates. Based on Irving AI Safety via Debate with Toulmin argumentation structure.

not rated 2 2mo ago A 37 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Flyvbjerg critical case methodology: select most-likely and least-likely cases to maximize inferential power.

not rated 2 2mo ago A 29 tokens original Apache-2.0

cross-examination

33

yogsoth-ai/stress-test

Skill Claude CodeCodex

Probes defender responses for inconsistencies, logical gaps, and unsupported claims. Acts as follow-up interrogation after initial defense.

not rated 2 2mo ago A 28 tokens original Apache-2.0

debate-architect

34

yogsoth-ai/stress-test

Skill Claude CodeCodex

Designs debate structure based on artifact type — selects attack vectors, assigns perspectives, determines escalation ladder, and configures round parameters.

not rated 2 2mo ago A 31 tokens original Apache-2.0

debate-critic

35

yogsoth-ai/stress-test

Skill Claude CodeCodex

Generates structured criticism from attack stance using Toulmin model. Produces claims, grounds, warrants, and rebuttals targeting artifact weaknesses.

not rated 2 2mo ago A 33 tokens original Apache-2.0

debate-defender

36

yogsoth-ai/stress-test

Skill Claude CodeCodex

Responds to attacks with counter-evidence and counter-arguments. Defends artifact using evidence, clarification, and rebuttal while acknowledging valid criticisms.

not rated 2 2mo ago A 34 tokens original Apache-2.0

debate-judge

37

yogsoth-ai/stress-test

Skill Claude CodeCodex

Evaluates debate exchanges, adjudicates argument quality, and produces round verdicts with confidence scores and reasoning.

not rated 2 2mo ago A 26 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Extracts key turning points, patterns, and insights from completed debate transcripts. Produces structured summary for verdict synthesis.

not rated 2 2mo ago A 29 tokens original Apache-2.0

deductive-chain

39

yogsoth-ai/stress-test

Skill Claude CodeCodex

Derive logical consequences step by step from a given premise, building a traceable derivation chain.

not rated 2 2mo ago A 24 tokens original Apache-2.0

design-fmea

40

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Research design-level FMEA — function analysis, failure mode identification, severity/occurrence/detection scoring per AIAG-VDA 2019.

not rated 2 2mo ago A 35 tokens original Apache-2.0

detection-scoring

41

yogsoth-ai/stress-test

Skill Claude CodeCodex

Rate detectability 1-10 (inverted: 10 = hardest to detect). Estimates how likely current controls would catch the failure before impact.

not rated 2 2mo ago A 35 tokens original Apache-2.0

devils-advocacy

42

yogsoth-ai/stress-test

Skill Claude CodeCodex

Construct the strongest possible counter-argument against a position, steelmanning the opposition before attacking.

not rated 2 2mo ago A 26 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Identifies agreement and disagreement patterns across multiple perspective evaluations. Maps consensus clusters and persistent divergence points.

not rated 2 2mo ago A 25 tokens original Apache-2.0

elegance-trap-probe

44

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Attack a beautiful unified result on the suspicion that its beauty is the bug. Distinguishes EARNED simplicity (forbids/predicts/subsumes) from DECORATIVE simplicity (re-describes/relabels/accommodates). Directly serves the Occam aesthetic by making it a falsifiable bar, not a vibe. Methods: Sober…

not rated 2 2mo ago A 103 tokens original Apache-2.0

evidence-scout

45

yogsoth-ai/stress-test

Skill Claude CodeCodex

Searches for external evidence supporting or opposing specific claims. Returns structured evidence with source assessment and relevance scoring.

not rated 2 2mo ago A 26 tokens original Apache-2.0

evidence-tournament

46

yogsoth-ai/stress-test

Skill Claude CodeCodex

Tactic: Evidence gathering, cross-examination, and quality judgment. External evidence is collected, presented, challenged, and scored for relevance and reliability.

not rated 2 2mo ago A 35 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Generate boundary and extreme test values for a given parameter dimension to stress-test claims.

not rated 2 2mo ago A 21 tokens original Apache-2.0

factor-enumeration

48

yogsoth-ai/stress-test

Skill Claude CodeCodex

List all key factors, conditions, and assumptions that support or enable the artifact's conclusion.

not rated 2 2mo ago A 23 tokens original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: