Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/xrensiu/claude-code-forge/adversarial-debuggingnpx skills add XRenSiu/claude-code-forge --skill adversarial-debugginggit clone --depth 1 https://github.com/XRenSiu/claude-code-forgeWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00102 | $0.03534 |
| Opus 5 | $0.00051 | $0.01767 |
| Sonnet 5 | $0.00020 | $0.00707 |
| Haiku 4.5 | $0.00010 | $0.00353 |
Grade A, and why
adversarial-debugging scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 422 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Adversarial Debugging
用多个 agent 的对抗辩论替代单 agent 的线性推理。
实测数据:复杂 bug 中单 agent 首次假设正确率约 40%。对抗式调试通过并行调查 + 相互挑战,将根因定位准确率提升到 80%+。
Announce at start: "I'm using the adversarial-debugging skill to create an agent team that investigates competing hypotheses in parallel."
前置条件: 需要启用 Agent Teams 实验性功能。 在 settings.json 中添加:
"env": { "CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS": "1" }
When to Use
digraph {
rankdir=TB;
start [label="遇到 Bug" shape=oval];
q1 [label="原因明显?" shape=diamond];
q2 [label="单 agent\n能定位?" shape=diamond];
q3 [label="多个可能\n根因?" shape=diamond];
systematic [label="systematic-debugging" shape=box];
adversarial [label="adversarial-debugging\n(本 Skill)" shape=box style=filled fillcolor=lightgreen];
direct [label="直接修复" shape=box];
start -> q1;
q1 -> direct [label="是"];
q1 -> q2 [label="否"];
q2 -> systematic [label="可能"];
q2 -> adversarial [label="困难"];
q3 -> adversarial [label="是"];
systematic -> q3 [label="失败后"];
}
vs. systematic-debugging
| 维度 | systematic-debugging | adversarial-debugging |
|---|---|---|
| Agent 数量 | 1 个 (顺序) | 3-7 个 (并行) |
| 假设处理 | 逐一测试 | 并行调查 + 辩论 |
| 偏见防御 | 流程纪律 | 结构化对抗 |
| 适合场景 | 常规 bug | 复杂/间歇性 bug |
| Token 消耗 | 低 | 高 (多 agent) |
| 速度 | 中等 | 快 (并行) |
| 准确率 | 高 | 更高 (多视角) |
The 5-Phase Protocol
┌─────────────────────────────────────────────────────────────────┐
│ ADVERSARIAL DEBUGGING │
├─────────────────────────────────────────────────────────────────┤
│ │
│ Phase 0: INTAKE 收集完整问题信息 │
│ ↓ │
│ │
│ Phase 1: HYPOTHESIZE 生成 3-5 个竞争假设 │
│ ↓ 每个假设必须可证伪、独立 │
│ │
│ Phase 2: TEAM ASSEMBLY 创建 Agent Team │
│ ↓ 为每个假设分配调查员 │
│ + Devil's Advocate + Synthesizer │
│ │
│ Phase 3: DEBATE 2-3 轮对抗辩论 │
│ ↓ 调查 → 挑战 → 回应 → 综合 │
│ │
│ Phase 4: VERDICT & FIX 共识判定 + TDD 修复 │
│ │
└─────────────────────────────────────────────────────────────────┘
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 422 lines · 102 tokens per session scan A a85f5c5e7f4e
adversarial-debugging is a skill published in the GitHub repository XRenSiu/claude-code-forge (2 stars, last pushed 1mo ago), licensed MIT. It adds 102 tokens to every session and 3,534 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
brainstorming
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.
auto-perf-optimize
Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.
chat-perf
Run chat perf benchmarks and memory leak checks against the local dev build or any published VS Code version. Use when investigating chat rendering regressions, validating perf-sensitive changes to chat UI, or checking for memory leaks in the chat response pipeline.
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…