Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add jackfranklin/dotfiles --skill adversarial-reviewergit clone --depth 1 https://github.com/jackfranklin/dotfilesWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/jackfranklin/dotfiles/adversarial-reviewer)<a href="https://agentmods.dev/skills/jackfranklin/dotfiles/adversarial-reviewer"><img src="https://agentmods.dev/badge/skills/jackfranklin/dotfiles/adversarial-reviewer.svg" alt="Measured on agentmods" height="20"></a>- NVIDIA SkillSpector warn
SkillSpector: 1 finding, up to high
These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →
- high Memory Poisoning · line 27 Skill manipulates agent memory, state, or stored context. Memory corruption can alter personality, override safety rules, or cause unpredictable behavior.Fix: Protect agent memory and state from modification by untrusted content. Use read-only memory for critical instructions and validate all state changes.
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00041 | $0.00487 |
| Opus 5 | $0.00020 | $0.00244 |
| Sonnet 5 | $0.00008 | $0.00097 |
| Haiku 4.5 | $0.00004 | $0.00049 |
Grade A, and why
adversarial-reviewer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Adversarial Audit
Overview
Perform a rigorous, adversarial review of proposed logic or changes to prevent happy-path bias. Identify security vulnerabilities, parsing anomalies, state desyncs, and boundary errors before writing code.
When to Use
Use when:
- Planning a code change or drafting an implementation plan.
- Reviewing code diffs for security, correctness, and edge cases.
- Writing parsers, serializers, UI components handling user input, or state transitions.
Core Pattern
Audit the changes against these four hazard vectors:
-
Input & Sanitization
- Control Characters: Are special syntactical characters (e.g.
*,_,\,`in markdown;<,>,&in HTML) escaped or sanitized? - Injection: Can scripts, event handlers, or harmful protocol schemes (
javascript:) be injected? - Encoding: How are Unicode characters, surrogate pairs, or invalid octets handled?
- Control Characters: Are special syntactical characters (e.g.
-
State & Concurrency
- UI & App Desync: Can fast user interactions (e.g. double clicks) trigger duplicate requests or corrupt state?
- Race Conditions: How does the system behave if async responses return out of order?
- Caching: Is stale cache cleared or invalidated?
-
Boundary Values
- Inputs: Handle null, undefined, empty strings, extremely large payloads, or deeply nested structures safely.
- Errors: Ensure timeouts, network dropouts, or permission rejections fail gracefully instead of crashing or leaking data.
-
Resource Lifecycle
- Leaks: Clean up active event listeners, timers, file handles, or network sockets when the component unmounts.
Common Mistakes
- Hacky regular expressions: Using naive regex for HTML/markdown escaping or sanitization instead of standard libraries/well-tested parsers.
- Silent failure: Swallowing errors without logging or notifying the user.
- Happy-path testing: Writing unit tests that only cover valid inputs.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 42 lines · 41 tokens per session scan A 4383f57559d0
adversarial-reviewer is a skill published in the GitHub repository jackfranklin/dotfiles (254 stars, last pushed 6d ago), licensed MIT. It adds 41 tokens to every session and 487 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
autoreview
Pre-commit/ship code review: Codex default; optional Claude or Pi.
rework-rate
Measure and interpret PR rework rate — the emerging 5th DORA metric.
omh-code-review
This is a Hermes-native code-review workflow skill.
revdiff-plan
Review the last Codex assistant message (plan, analysis, or proposal) with inline annotations in a TUI overlay. Extracts the most recent response from Codex rollout files and opens it in revdiff for review and annotation. Activates on "revdiff-plan", "review plan with revdiff", "annotate plan", "review last response"…
code-reviewer
Code review specialist focused on patterns, bugs, security, and performance.
agent-teams-simplify-and-harden
Implementation + audit loop using parallel agent teams with structured simplify, harden, and document passes. Spawns implementation agents to do the work, then audit agents to find complexity, security gaps, and spec deviations, then loops until code compiles cleanly, all tests pass, and auditors find zero issues or…