challenge

challenge is a command for coding agents from s977043/river-review. It costs 23 tokens per session (586 once invoked), scanned A, original, MIT.

An adversarial review command that examines current code changes from three angles: imagining future failures, simulating possible abuse, and testing the logic behind design decisions. It links findings to specific changed lines.

In plain words
What is it for?
Use it to review a diff, staged changes, a file, directory, or pull request for pre-mortem risks, attack paths, questionable reasoning, and missing safeguards.
Why use it?
It helps uncover security gaps, weak assumptions, and failure scenarios before a change is merged. Each reported finding includes a concrete next action.

Command

Part of the river-review plugin — 140 skills, 15 commands, 5 agents, 2 hooks shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add commands/s977043/river-review/challenge
Clone the repo
git clone --depth 1 https://github.com/s977043/river-review

Or install river-review, the plugin that ships this one along with the rest of its 140 skills, 15 commands, 5 agents, 2 hooks.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for challenge

README.md
[![agentmods](https://agentmods.dev/badge/commands/s977043/river-review/challenge.svg)](https://agentmods.dev/commands/s977043/river-review/challenge)
Your own site
<a href="https://agentmods.dev/commands/s977043/river-review/challenge"><img src="https://agentmods.dev/badge/commands/s977043/river-review/challenge.svg" alt="Measured on agentmods" height="20"></a>
Per session 23 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 586 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00023 $0.00586
Opus 5 $0.00012 $0.00293
Sonnet 5 $0.00005 $0.00117
Haiku 4.5 $0.00002 $0.00059

Measured 5d ago against content hash 4f7d8c350948, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-05, from the pricing page.

Security

Grade A, and why

challenge scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

commands/challenge.md · 69 lines

What it actually says

Context

  • Status: git status
  • Diff: git diff
  • Recent commits: git log --oneline -10

Task

あなたは 敵対的レビュアー です。変更を3つの視点から徹底的に検証してください。

手法1: Pre-mortem(失敗シナリオ分析)

「この変更が6ヶ月後にインシデントを引き起こした」と仮定し、その原因を逆算せよ。

  • 崩れる前提は何か?
  • 因果連鎖はどうなるか?
  • 事前に防ぐ/検知する方法は?

手法2: War Game(攻撃者シミュレーション)

攻撃者の立場で、この変更をどう悪用できるかを分析せよ。

  • 新たに露出する攻撃面は?
  • 具体的な攻撃手順は?
  • 防御のギャップと最小限の対策は?

手法3: Logic Torturing(論理検証)

変更に含まれる設計判断の論理的な穴を突け。

  • この判断の前提は常に成立するか?
  • 代替案はなぜ棄却されたか?
  • この判断が間違いだとわかったとき、元に戻せるか?

Output Format

## 🔍 Adversarial Review

### Pre-mortem(失敗シナリオ)

<file>:<line>: [シナリオ] ...

### War Game(攻撃シナリオ)

<file>:<line>: [シナリオ] ...

### Logic Torturing(論理検証)

<file>:<line>: [検証] ...

### 最も重大な発見

<1件の要約と推奨アクション>

Rules

  • 各手法で最大3件(合計最大9件)に絞る(SKILL.md フル実行時は最大5件)。
  • すべての指摘は差分の具体的な行に紐づける。
  • 推測は推測として明示する。
  • 指摘には必ず次のアクション(Fix)を添える。
  • スキル定義の詳細: skills/agent-skills/adversarial-review/SKILL.md
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 69 lines · 23 tokens per session scan A 4f7d8c350948

Subscribe to this mod's changes

challenge is a command published in the GitHub repository s977043/river-review (3 stars, last pushed today), licensed MIT. It adds 23 tokens to every session and 586 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.