yogsoth-ai/stress-test

Research Artifact Stress-Testing Engine — five-campaign adversarial validation producing weakness-annotated verification reports

2Stars on the repository
110Mods indexed here, across every type
2mo agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

yogsoth-ai/stress-test

Skill Claude CodeCodex

Evaluate the probability of necessity (PN) for a causal factor — would the conclusion fail if this factor were absent?

not rated 2 2mo ago A 28 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Probability of Necessity and Sufficiency (PNS/PS) — systematically evaluate whether each factor is necessary, sufficient, both, or neither for the conclusion.

not rated 2 2mo ago A 40 tokens original Apache-2.0

occurrence-scoring

75

yogsoth-ai/stress-test

Skill Claude CodeCodex

Rate failure mode occurrence probability 1-10. Estimates how likely each failure mode is to manifest during research execution.

not rated 2 2mo ago A 28 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Build a detailed adversarial persona with background, motivation, expertise, blind spots, and preferred attack patterns.

not rated 2 2mo ago A 25 tokens original Apache-2.0

perspective-critic

78

yogsoth-ai/stress-test

Skill Claude CodeCodex

Evaluates artifact from a specific assigned perspective. Produces assessment grounded in that viewpoint's values, priorities, and expertise.

not rated 2 2mo ago A 29 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Execute Klein pre-mortem protocol — assume failure has occurred, generate plausible failure scenarios through prospective hindsight.

not rated 2 2mo ago A 29 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Tactic: Pre-mortem rapid screening feeds high-risk items into full FMEA analysis. Bridges fast intuitive generation with systematic structured analysis.

not rated 2 2mo ago A 37 tokens original Apache-2.0

probe-execution

81

yogsoth-ai/stress-test

Skill Claude CodeCodex

Execute a single attack probe against an artifact, record the result with evidence and severity classification.

not rated 2 2mo ago A 22 tokens original Apache-2.0

process-fmea

82

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Research execution process FMEA — analyzes how the research process itself can fail during execution, distinct from design-level failures.

not rated 2 2mo ago A 29 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Klein pre-mortem — assume the artifact has failed, then retrospect plausible causes. Generates rapid failure scenario catalog.

not rated 2 2mo ago A 30 tokens original Apache-2.0

re-scoring

84

yogsoth-ai/stress-test

Skill Claude CodeCodex

Re-evaluate S/O/D scores after mitigation measures are in place. Validates that mitigations actually reduce risk as expected.

not rated 2 2mo ago A 29 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Systematic adversarial probing retuned for truth-seeking. Threat surface = the set of load-bearing claims. Output is NOT a resilience score and NOT a hardening list — it is, per claim, the specific observation/computation that would refute it, plus which attacks succeeded. Methods: UFMCS Key Assumptions…

not rated 2 2mo ago A 91 tokens original Apache-2.0

red-teaming

86

yogsoth-ai/stress-test

Skill Claude CodeCodex

Campaign: Systematic adversarial attack from military/intelligence/AI-safety traditions. Core question: Can systematic adversarial attacks find fatal flaws? Methods: UFMCS Red Team Handbook v9.0, CIA SAT, Anthropic Red Teaming, NIST AI RMF, Inie et al. 12-strategy taxonomy.

not rated 2 2mo ago A 72 tokens original Apache-2.0

risk-prioritization

87

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Action Priority matrix — classifies failure modes into H/M/L priority using severity-weighted scoring per AIAG-VDA 2019 Action Priority tables.

not rated 2 2mo ago A 38 tokens original Apache-2.0

severity-scoring

88

yogsoth-ai/stress-test

Skill Claude CodeCodex

Rate failure mode severity 1-10 based on end-effect impact. Follows AIAG-VDA severity scale calibrated for research artifacts.

not rated 2 2mo ago A 31 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Remove one specified factor from the artifact's support structure and reason about how the conclusion changes.

not rated 2 2mo ago A 23 tokens original Apache-2.0

society-of-mind

90

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Multi-agent collaborative debate based on Du et al. Society of Mind. Agents share perspectives iteratively until convergence or divergence is detected.

not rated 2 2mo ago A 34 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Military-grade assumption testing — Key Assumptions Check, Devil's Advocacy, Team A/B analysis to expose hidden dependencies and unexamined beliefs.

not rated 2 2mo ago A 38 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Tactic: Progressive debate escalation based on confidence thresholds. Each round increases attack sophistication until defender collapses or proves resilient.

not rated 2 2mo ago A 33 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Paper landscape scan returning abstracts and metadata. Import of literature-engine/paper-overview skill. Abstracts only — no conclusions from abstracts.

not rated 2 2mo ago A 33 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Paper full text access via alphaxiv answerpdfqueries or getpapercontent(fullText=true). Import of literature-engine/paper-research skill. Raw extracted text for precise claims.

not rated 2 2mo ago A 43 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Paper AI summary report via alphaxiv getpapercontent. Import of literature-engine/paper-search skill. Structured AI-generated intermediate report.

not rated 2 2mo ago A 33 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Tactic: Sequential perspective evaluation with divergence aggregation. Each agent evaluates from a distinct viewpoint, then disagreements are surfaced and resolved.

not rated 2 2mo ago A 32 tokens original Apache-2.0

At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: