yogsoth-ai/stress-test

Research Artifact Stress-Testing Engine — five-campaign adversarial validation producing weakness-annotated verification reports

2Stars on the repository
110Mods indexed here, across every type
2mo agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

yogsoth-ai/stress-test

Skill Claude CodeCodex

Determines whether validation has reached saturation — no new weaknesses or failure modes being discovered. Used by all 5 campaigns as termination signal.

not rated 2 2mo ago A 34 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Deep web full-text retrieval via Apify RAG browser. Import of web-browsing/web-research skill. Full page content for substantive analysis.

not rated 2 2mo ago A 36 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Quick web scanning for landscape understanding. Import of web-browsing/web-search skill. Snippets only — no conclusions from snippets alone.

not rated 2 2mo ago A 32 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Pearl Three-Step counterfactual — Abduction (fit model to evidence), Action (intervene on factor), Prediction (derive counterfactual outcome).

not rated 2 2mo ago A 38 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Tactic: Full attack lifecycle — threat surface enumeration, attack vector generation, systematic probing, and finding aggregation across all surfaces.

not rated 2 2mo ago A 31 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Evaluate the probability of sufficiency (PS) for a causal factor — would this factor alone be enough to produce the conclusion?

not rated 2 2mo ago A 31 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Tactic: List all factors, remove one at a time, assess conclusion stability, rank factors by load-bearing importance.

not rated 2 2mo ago A 30 tokens original Apache-2.0

systematic-probing

106

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: AI-safety systematic probing — enumerate all threat surfaces, generate attack vectors per surface, execute probes, and aggregate findings across the full attack space.

not rated 2 2mo ago A 36 tokens original Apache-2.0

thought-experiment

107

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Williamson-style precise thought experiments — construct carefully specified counterfactual scenarios to test whether conclusions depend on contingent features.

not rated 2 2mo ago A 28 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Enumerate all attackable surfaces of an artifact — logical, empirical, methodological, social, and practical dimensions.

not rated 2 2mo ago A 29 tokens original Apache-2.0

verdict-synthesis

109

yogsoth-ai/stress-test

Skill Claude CodeCodex

Synthesizes findings from a completed campaign into typed verdict reports. Produces DebateVerdict, RedTeamReport, FailureAnticipationReport, CounterfactualMap, or AdversarialStressReport depending on campaign. Also supports cross-campaign StressTestSummary.

not rated 2 2mo ago A 58 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Classifies discovered weaknesses into severity tiers (fatal/major/minor/cosmetic) with structured justification and exploitability assessment.

not rated 2 2mo ago A 30 tokens original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: