Getting it into your agent
It runs from inside its repository, so the clone comes first — what it calls does not travel with the file alone.
git clone --depth 1 https://github.com/realnaka/alphaloopnpx agentmods add skills/realnaka/alphaloop/agent-tool-escalationWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/realnaka/alphaloop/agent-tool-escalation)<a href="https://agentmods.dev/skills/realnaka/alphaloop/agent-tool-escalation"><img src="https://agentmods.dev/badge/skills/realnaka/alphaloop/agent-tool-escalation/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/realnaka/alphaloop/agent-tool-escalation"><img src="https://agentmods.dev/badge/skills/realnaka/alphaloop/agent-tool-escalation.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00117 | $0.07075 |
| Opus 5 | $0.00059 | $0.03537 |
| Sonnet 5 | $0.00023 | $0.01415 |
| Haiku 4.5 | $0.00012 | $0.00707 |
Grade C, and why
agent-tool-escalation scanned grade C with 2 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Harvests environment variableshighData exfiltration
Enumerating or grepping the environment for keys collects credentials unrelated to what the mod says it does.
env | grep -i KEY_NAME Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
curl -s "https://qt.gtimg.cn/q=hk00877" How it starts
The opening of the file, as written. The whole thing — 397 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Agent Tool Escalation
核心原则:低权限工具失败时升级工具,不要升级用户。 想要硬保障的话,把本文「自检三问」沉淀成 agent 的 always-apply 规则(各家 agent 的规则机制见
../openorder/docs/compatibility.md);本 skill 是「按需展开」的完整参考。
AskQuestion 之前的自检三问(每次失败都跑)
- 还有什么工具我没试?
- 有没有一条 ≤3 秒的环境探针能澄清这个不确定性?
- 如果用户告诉我答案,他会反问"你为什么不自己查"吗?
三问全过 → 才允许 AskQuestion。
「修正已有信息」之前的自检三问(NEW,Case 3 教训)
任何"我发现 v1 写错了,要修正"的瞬间,先停下,跑这三问:
- 我打算推翻的是什么 —— 是新发现 vs 已有 baseline 的冲突,还是仅凭"我以为"的内省怀疑?
- 支持我修正的证据是什么 —— 是官方公告 / 一手数据 / 用户原话(ground truth),还是逻辑推断 / 类比 / "印象"?
- 如果我错了,污染半径多大 —— 这个修正会衍生几个论断?会被几个文件引用?会污染哪个 framework?
任何一问不通过 → 必须先 WebFetch 官方源 / 跑环境探针 / 直接 AskQuestion,不允许直接修正。
核心原则:补充新信息的证据门槛 < 修正已有信息的证据门槛 < 全面推翻已有结论的证据门槛。门槛比例 1 : 10 : 100。
Case 6:vector 路径搜索反模式(NEW,2026-05-10 晚)
反模式描述
判断"A 跟 B 没有关系 / 没有合作 / 是 narrative 不是事实"时,只用 1-2 条直接关键词搜索(如 "A B partnership"、"A B Inphi"),没有穷尽 vector 路径:
- A → 第三方 → B(A 卖材料给 C,C 被 B 收购)
- A → 共同技术 → B(A、B 都用同一种技术 / 平台 / foundry)
- A → 共同投资人 → B(A、B 都被同一基金投)
- A → 共同客户 → B(A、B 都卖给同一家 hyperscaler)
- A → 论文/学术 → B(A 在 B 的论文 acknowledgment 里被点名)
结果:在没找到直接合作公告时,错误地下"没有合作"结论 + 强烈用词("散户 narrative"、"不是事实"),并把这个错误结论沉淀进知识库。
真实案例(2026-05-10 LWLG-Polariton-MRVL)
| 阶段 | 行为 | 错误本质 |
|---|---|---|
| 17:25 v1.0 | 用户问 "LWLG 是不是有个技术和 MRVL 密切相关?" | 触发 |
| 17:25 v1.0 | 我搜 "Lightwave Logic Marvell partnership" / "LWLG Marvell Inphi" → 找不到 | 只用直接关键词 |
| 17:25 v1.0 | 写入 LWLG.md v1.0:「LWLG 跟 MRVL 没有任何官方公告的直接合作」+「投资者论坛叙事是散户 narrative,不是事实」 | 强烈否定 + 沉淀进档案 |
| 17:35 | 用户提供 PhotonCap "What Marvell Bought Was the Slot" + 2 张截图 | ground truth 触发 |
| 17:50 v1.1 | 5 个独立官方源验证:MRVL 4/22 收购 Polariton + Polariton 5+ 年用 LWLG Perkinamine + Optica 2025 record device 论文 acknowledgment 明确名 LWLG | 完全推翻 v1.0 |
→ 正确搜索路径:LWLG Polariton + MRVL Polariton + Polariton ETH spinoff + Polariton plasmonic chromophore —— 任何一条都能立即命中 5+ 年合作历史 + 4/22 收购公告。
「下"没有 X"结论」之前的自检三问
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 397 lines · 117 tokens per session scan C dcdf7d480b48
agent-tool-escalation is a skill published in the GitHub repository realnaka/alphaloop (19 stars, last pushed 2mo ago), licensed MIT. It adds 117 tokens to every session and 7,075 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it C with 2 findings (harvests environment variables, makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
roblox-performance
Use when profiling Roblox performance or diagnosing FPS, memory, network, mobile, or hot-path problems.
incident-responder
Structured incident response workflow — detect, investigate, communicate, and resolve using session data as evidence. Use when an incident is declared, a production issue is reported, or the on-call engineer needs session evidence.
error-forensics
Investigate JavaScript errors, console errors, network failures, and crashes in production. Use when debugging user-reported issues, tracking error frequency, finding sessions with specific errors, or determining whether a bug is widespread or isolated.
session-review
Use when diagnosing user-reported issues, investigating bugs, analyzing user behavior, or validating UI correctness using Fullstory session recordings.
apple-cleanup
Exhaustive engineering hardening of an iOS app. Reviews for Swift 6 compliance, crash risks, App Store rejection risks, and tech debt; builds a surgical plan; dispatches parallel subagents to fix all P0-P2 issues; then pushes an alpha to TestFlight. Use for pre-submission cleanup and code hardening, not design polish.
apple-polish
Design and keynote-readiness craftsmanship review of an iOS app. Evaluates through Jony Ive (visual obsession) and Steve Jobs (demo readiness) perspectives, presents prioritized findings, then orchestrates parallel agents to fix selected issues and push a TestFlight build. Use for design polish, not engineering bugs.