adversarial testing skills

80 tagged adversarial testing, measured the same way as everything else here.

Browse within: Multi-Agent 61research-validation 60stress-testing 60ai-code-generation 7ai-security 7devsecops 7llm-security 7ideation 5

darwinia

01

0xSanei/darwinia

Skill Claude CodeCodex

Evolve trading strategies through genetic algorithms and adversarial combat. No API keys, no cloud — pure Python + numpy.

86 4mo ago A 28 tokens original MIT

darwinia

02

0xSanei/darwinia

Skill Claude CodeCodex

Evolve trading strategies through genetic algorithms and adversarial combat. Run Darwinian selection on BTC data to discover battle-tested strategies. No API keys, no cloud.

86 4mo ago A 36 tokens original MIT

crucible

03

smshahbaj/crucible

Skill Claude CodeCodex

Pressure-test important decisions, plans, proposals, strategies, architectures, and recommendations before acting. Use when trade-offs, uncertainty, meaningful downside, or an important second opinion could change the outcome. Stay lightweight for routine or trivial work.

9 8d ago A 50 tokens original MIT

devils-advocate

04

alejandrosaenz117/bonfires-marketplace

Skill Claude CodeCodex

This skill should be used when the user asks for "an adversarial review", "security review", "devil's advocate", "what could go wrong", "find the vulnerability", "threat analysis", "penetration test this", or "challenge this design". Use this skill when the user wants to identify security threats and architectural…

4 2d ago A 80 tokens original MIT

alejandrosaenz117/bonfires-marketplace

Skill Claude CodeCodex

Prioritize any workload using the Eisenhower Matrix. Use this skill whenever a user provides a brain dump of work—Jira summaries, meeting notes, sprint dumps, task lists, or prose descriptions of their workload. This skill categorizes tasks into four quadrants: Q1 (Urgent+Important: do now), Q2 (Important+NotUrgent…

4 2d ago A 174 tokens original MIT

osv-scanner

06

alejandrosaenz117/bonfires-marketplace

Skill Claude CodeCodex

This skill should be triggered when the user asks about dependency security, vulnerability scanning, or package safety. Examples: "check my dependencies for vulnerabilities", "scan my packages", "are my dependencies safe", "dependency audit", "check for CVEs", "security audit", "vulnerable packages", "scan…

4 2d ago A 71 tokens original MIT

food-chain-code

07

CodedRichy/food-chain-ideation

Skill Claude CodeCodex

Adversarial architecture stress-tester. Selects attacker agents from a code-specific behavioral DNA library matched to the technical decision. Each attacks under strict role-lock. Weakest eliminated, survivor absorbs and evolves. Tests architecture decisions before a line of code is written. Works in Claude.ai, Claude…

3 2mo ago A 75 tokens original MIT

food-chain-ideation

08

CodedRichy/food-chain-ideation

Skill Claude CodeCodex

10-minute pre-ship stress-test. Run this before you ship, not instead of shipping. Dynamically selects animal agents from a behavioral DNA library matched to the specific problem, each attacking under strict role-lock with zero shared context. Weakest eliminated each round, survivor absorbs the insight and evolves…

3 2mo ago A 120 tokens original MIT

food-chain-pitch

09

CodedRichy/food-chain-ideation

Skill Claude CodeCodex

Investor pitch stress-tester. Spawns adversarial agents from a pitch-specific behavioral DNA library. Each attacks the pitch narrative, financial model, or due diligence surface under strict role-lock. Elimination rounds harden the pitch. Works in Claude.ai, Claude Code, Cursor, Windsurf, Copilot with no dependencies.

3 2mo ago A 69 tokens original MIT

adversary

10

tasumermaf/the-adversary

Skill Claude CodeCodex

Run The Adversary v2 — the three-stage adversarial audit (CHECK deterministic, FIND independent lenses, VERIFY reproduce-or-demote) over an artifact at a pinned commit. Use when the user says "run the adversary", "adversarial audit", "audit this paper/repo", "release gate", or "gate this before I ship". Picks the…

2 28d ago A 118 tokens original MPL-2.0

closest-worlds

11

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Lewis Possible Worlds — find the minimal change to reality that would flip the conclusion, measuring how close the nearest world where the conclusion fails.

2 2mo ago A 33 tokens original Apache-2.0

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Legal adversarial structure — prosecution presents case, defense responds, evidence is cross-examined, judge delivers verdict. Emphasizes evidence quality and procedural rigor.

2 2mo ago A 39 tokens original Apache-2.0

design-fmea

13

yogsoth-ai/stress-test

Skill Claude CodeCodex

Strategy: Research design-level FMEA — function analysis, failure mode identification, severity/occurrence/detection scoring per AIAG-VDA 2019.

2 2mo ago A 35 tokens original Apache-2.0

gauntlex:doctor

14

sanjoy1234/gauntlex

Skill Claude CodeCodex

Full GAUNTLEX environment health check — model reachable, ChromaDB writable, AVF fixtures pass, no unexpected outbound network (air-gap verification). Use before first run or when debugging failures.

1 1mo ago A 45 tokens original MIT

gauntlex:run

15

sanjoy1234/gauntlex

Skill Claude CodeCodex

Run a full GAUNTLEX adversarial session on a spec file or GitHub issue URL. Invokes the Gauntlex (Builder + Breaker concurrent) and outputs an Adversarial Resilience Score (ARS) with a Resilience Report.

1 1mo ago A 56 tokens original MIT

gauntlex:validate

16

sanjoy1234/gauntlex

Skill Claude CodeCodex

Dry-run GAUNTLEX validation — parses the spec, resolves policy playbooks, checks model availability, and runs the AVF golden fixture gate. No attacks are fired. Exit 1 if any check fails.

1 1mo ago A 48 tokens original MIT

tournament-judge

17

alextverdyy/tournament-judge

Skill Claude CodeCodex

Runs a domain-neutral, evidence-based tournament in which independent candidates are reviewed, anonymized, and scored by a blind judge against a rubric fixed before judging. Use when comparing competing plans, designs, implementations, documents, tools, vendors, strategies, or other substantial options; when the user…

1 23d ago A 105 tokens original MIT