devil

A reasoning review tool that argues against the current interpretation, plan, or decision. It looks for the assumption that the conclusion depends on most and examines that point in depth.

In plain words
What is it for?
Use it to challenge hypotheses, plans, interpretations, and irreversible decisions by finding their key assumption and developing counter-arguments.
Why use it?
It helps expose weak reasoning before a plan is acted on, especially when a decision may be difficult to undo. It reviews the thinking itself rather than checking code or files.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/jun0-ds/sonmat/devil
Any agent
npx skills add jun0-ds/sonmat --skill devil
Clone the repo
git clone --depth 1 https://github.com/jun0-ds/sonmat

Made for: Claude Code, Codex.

Per session 48 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 3,454 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00048 $0.03454
Opus 5 $0.00024 $0.01727
Sonnet 5 $0.00010 $0.00691
Haiku 4.5 $0.00005 $0.00345

Measured 2d ago against content hash 21885c7d1cc0, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

devil scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/devil/SKILL.md · 221 lines

How it starts

The opening of the file, as written. The whole thing — 221 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Devil — Devil's Advocate for Thinking

Discovery-led counter-argument against the current interpretation, hypothesis, or plan: locate the single load-bearing assumption first, then drive depth into the one axis where the claim is actually thin. Targets reasoning and judgment, not code or artifacts.

Positioning within sonmat's verification axes:

Skill Axis Asks
guard Main-side verification "Is the work operationally safe?"
inspect System impact "What could break?"
witness Intent-artifact match "Does this match what the user asked for?"
punch Completeness + residue "Is anything missing or left over?"
scribe Post-work persistence "Is anything worth keeping?"
devil Reasoning "Is the thinking itself sound?"

Invoke: /devil (runs once for the current claim, delivers a balance table, and exits — not a mode). The user can re-invoke for another round.


What devil does

When invoked, take the current interpretation and find its load-bearing part first, then drive depth there. Discovery before depth; asymmetric attack after discovery. The steps below are not a broad-front assault — they are a structured way to locate the one place the claim is actually fragile and pressure-test it there.

1. Identify the claim

Extract the core claim(s) being made. State them back clearly so the user can confirm what's being challenged.

[devil] Challenging: "{the claim}"

2. Active discovery — find the load-bearing assumption

Before attacking on any axis, do the discovery-led step first: find the single thing this claim is standing on. Depth is not a dial to turn up at the start; depth is what follows naturally once you have located the load-bearing part.

Use devil CCT as the discovery checklist (analogous to chess's CCT — Checks / Captures / Threats — which is a compressed triage that surfaces the sharp part of a position before any deep calculation):

Check Question What it surfaces
Claim-crux What is the one thing that, if false, would flip this claim? The load-bearing assumption
Counter-fit Does the same evidence also fit an opposite conclusion? If so, what distinguishes them? A hidden alternative riding the same data
Cause-chain Is the cause → effect direction actually established, or only correlated / reversed / mediated? A reversed or spurious causal link

Read the full file on GitHub · 221 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 221 lines · 48 tokens per session scan A 21885c7d1cc0

Subscribe to this mod's changes

devil is a skill published in the GitHub repository jun0-ds/sonmat (6 stars, last pushed 4d ago), licensed BSD-3-Clause. It adds 48 tokens to every session and 3,454 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

mundo

MUNDO - THE EMPEROR. The ultimate AI learning engine. No triggers needed. Every task, Mundo takes over. Consults ALL AIs, crawls ALL web, integrates ALL solutions, saves ALL useful skills. Self-evolving. Collective consciousness. Infinite growth. Uses Three Departments and Six Ministries system to rule all skills.…

LiHongwei-cn/lihongwei-cn · 91 tokens

llm-app-patterns

Production-ready patterns for building LLM applications, inspired by Dify and industry best practices.

LiHongwei-cn/lihongwei-cn · 36 tokens

kpi-dashboard-design

Comprehensive patterns for designing effective Key Performance Indicator (KPI) dashboards that drive business decisions.

LiHongwei-cn/lihongwei-cn · 24 tokens

neat-freak

End-of-session knowledge cleanup with OCD-level rigor — reconciles project docs (CLAUDE.md, README.md, docs/) and agent memory against the code so nothing rots. 会话结束后对项目文档和记忆进行洁癖级审查与同步。MUST trigger when the user says: "sync up", "tidy up docs", "update memory", "clean up docs", "/sync", "/neat", "同步一下", "整理文档"…

LiHongwei-cn/lihongwei-cn · 215 tokens

nature-reader

Build full-paper Chinese-English side-by-side, figure/table-aware, source-grounded Markdown readers for journal or conference papers from PDF, DOI, arXiv, publisher HTML, or pasted text. Use whenever the user asks to translate or read a paper, make 中英文对照/原文对照/全文翻译解读, extract figures or tables into the right positions…

LiHongwei-cn/lihongwei-cn · 116 tokens

resume-builder

制作/优化专业简历(HTML格式)。触发词:简历、resume、CV、求职、找工作、投简历。自动加载,白底专业风格,禁止深色主题和过度设计。.

LiHongwei-cn/lihongwei-cn · 52 tokens