Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add bahayonghang/my-ai-cli-toolkit --skill dual-steelmangit clone --depth 1 https://github.com/bahayonghang/my-ai-cli-toolkitWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/bahayonghang/my-ai-cli-toolkit/dual-steelman)<a href="https://agentmods.dev/skills/bahayonghang/my-ai-cli-toolkit/dual-steelman"><img src="https://agentmods.dev/badge/skills/bahayonghang/my-ai-cli-toolkit/dual-steelman/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/bahayonghang/my-ai-cli-toolkit/dual-steelman"><img src="https://agentmods.dev/badge/skills/bahayonghang/my-ai-cli-toolkit/dual-steelman.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00195 | $0.02210 |
| Opus 5 | $0.00097 | $0.01105 |
| Sonnet 5 | $0.00039 | $0.00442 |
| Haiku 4.5 | $0.00019 | $0.00221 |
Grade A, and why
dual-steelman scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 126 lines — stays where its author put it; the contents beside it link to each section on GitHub.
dual-steelman: 双向钢人论证
把"直接给答案"换成"先把双方都武装到最强,再逼出一个明确判断"。
这个 skill 对抗两件事:模型顺着用户说(谄媚),以及用户嘴上问的不是真正想解决的问题。角色分工是固定的:你是执行钢人论证的人;用户的立场是第一个被强化的对象;反对用户的立场是第二个被强化的对象。 你不是辩手,也不是和事佬。
Usage
Instructions
0. 适用判断
- 输入必须是一个待决的判断、选择或立场。事实性提问、代码任务、要求直接执行的事务,不走本流程,按正常方式回答。
- 用户只说"用钢人论证/帮我想清楚"但没给具体问题:先请用户给出问题,不要空转。
- 用户明确说"直接给结论,不要流程":压缩执行——每方钢人各 3 句以内、跳过停顿、当轮给判断,并注明这是压缩版。
第 1-4 步在同一轮回复里完成,第 4 步末尾停住;第 5 步在用户回答后单独进行。
1. 重述真实问题
- 用最完整、最有力的方式重述用户真正想解决的问题,而不是字面问题。表面问题和真实关切经常分离:问"要不要辞职"的人,真实关切可能是"三年后我还有没有竞争力"。把你识别到的真实关切明写出来。
- 问题复合时拆成 2-5 个子问题,标出哪一个是主问题。
- 重述必须使用用户给出的具体事实(数字、日期、金额、人名、约束)。禁止抽象成"某公司""某个选择"。
- 结尾加一句"如果重述有偏差,请直接纠正",然后继续往下走,不在这里停。
2. 双向钢人
按输入形态选择结构:
- 用户有明确立场(要不要/该不该):正方 = 用户当前想法的最强版本;反方 = 反对它的最强版本。
- 多个候选项(选 A/B/C):对每个候选分别给出"支持它的最强论证 + 反对它的最强质疑"。候选超过 4 个时,先按用户约束收缩到 2-4 个再钢人化,并说明淘汰理由。
- 用户没有立场:从重述中提炼 2-3 个可辩立场,再双向钢人。
钢人质量标准(本 skill 的核心,逐条执行):
- 每一方的论证要强到该立场最聪明的支持者愿意签名认领。自检问句:"持这个立场的人会承认这是公平的表述吗?"不会,就继续加强。
- 用该方能拿出的最好证据和最合理的价值排序,不是模板化优缺点清单。
- 反方必须包含"正方最难回答的那个质疑";正方必须包含"反方最难反驳的那个理由"。
- 每方论证都落在用户的具体事实上。
- 禁止:稻草人化任何一方;两边写成对称的客套话;用篇幅或措辞偏袒一方;在这一步泄露你的倾向。
3. 真实分歧与关键变量
- 用一句话点名双方真正的分歧,并标出类型:价值排序分歧 / 事实预测分歧 / 概念定义分歧。
- 找出 1-2 个最可能改变结论的关键变量,用条件句写出翻转关系:"若 X 成立,倾向 A;若 Y 成立,倾向 B。"
- 变量必须是用户可回答或可查证的,不允许写成"取决于你自己"。
4. 只问一个问题(硬性停顿)
- 从关键变量中挑信息量最大的一个问题。常用形态:十年后回望("十年后回头看,你更希望当时……")、价值排序二选一、具体量化阈值。
- 只问这一个。不给问题清单,不三连问,不附答案选项分析。
- 问完立即结束本轮回复。不要在同一轮给出判断、倾向、或"如果你问我的话"式暗示。 这是本 skill 最容易失守的一步。
- 唯一例外:用户最初的提问里已经明确回答了这个关键变量。此时引用用户原话说明为何不需停顿,直接进入第 5 步。
5. 用户回答后:判断、理由、行动
- 明确判断:一句话站定一边(或站定某个候选)。禁止"两边都有道理""看情况"式收尾。
- 理由:引用钢人论证中真正决定性的 1-3 条,说明在用户给出的回答下,为什么另一方的最强论证仍然不足以翻盘。
- 下一步行动:2-4 条,具体、可开始、有先后。
- 用户拒绝回答或回答"都想要":按关键变量给条件式判断("若 X 则 A;若 Y 则 B"),每个分支仍是站定的结论,不给和稀泥总结。
- 判断必须对这个用户有用:引用其具体事实和第 4 步的回答,不输出谁都能套用的通用建议。
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 126 lines · 195 tokens per session scan A 79d5a3e8eec8
dual-steelman is a skill published in the GitHub repository bahayonghang/my-ai-cli-toolkit (16 stars, last pushed yesterday), licensed MIT. It adds 195 tokens to every session and 2,210 once invoked, about $0.0010 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
clarify
Decides whether to ask clarifying questions or proceed with an answer, optimizing for information value vs. delay cost.
devils-advocate
Utiliser avant toute décision importante pour faire jouer à l'IA le rôle d'un sceptique professionnel. Produit les trois meilleures raisons de ne pas faire ce qui est proposé, classées par force décroissante.
decision-heuristics
A set of heuristics for making difficult personal decisions such as changing jobs, buying a home, moving, forming a partnership, or getting married. It is intended for major choices, not everyday decisions.
acceptance
A personal decision framework for difficult situations, built around three choices: change the situation, accept it, or leave it.
exec-briefing-memo
A one-page decision memo that condenses complex material into the problem, recommended choice, evidence, trade-offs, risks, and next actions. It is meant for a decision maker, not as meeting notes, a weekly report, or a product requirements document.
thinking-model-router
When unsure which thinking skill fits, map domain and problem type, then return NONE or one primary skill by default (at most three complementary).