Getting it into your agent
It runs from inside its repository, so the clone comes first — what it calls does not travel with the file alone.
git clone --depth 1 https://github.com/realnaka/alphaloopnpx agentmods add skills/realnaka/alphaloop/claim-verificationWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/realnaka/alphaloop/claim-verification)<a href="https://agentmods.dev/skills/realnaka/alphaloop/claim-verification"><img src="https://agentmods.dev/badge/skills/realnaka/alphaloop/claim-verification/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/realnaka/alphaloop/claim-verification"><img src="https://agentmods.dev/badge/skills/realnaka/alphaloop/claim-verification.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00162 | $0.02036 |
| Opus 5 | $0.00081 | $0.01018 |
| Sonnet 5 | $0.00032 | $0.00407 |
| Haiku 4.5 | $0.00016 | $0.00204 |
Grade A, and why
claim-verification scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 85 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Claim Verification(二手信息核实方法论)
核心原则:任何二手信息(推文 / 小作文 / 新闻 / 研报 / 其他 AI 的结论)默认是假设(hypothesis),不是事实。验真前不引用、不写入知识库、不驱动决策。
姊妹 skill:
agent-tool-escalation管"工具怎么用";本 skill 管"信息怎么验真"。投资场景配合openorder落档。
5 步工作流
1. 拆解(DECOMPOSE) 把长 thesis / 一段话拆成「原子声明」,每条只含一个可独立验证的事实
2. 定级(TRIAGE) 给每条声明标来源可信度 + 优先级(影响大 + 易错 = 先查)
3. 查源(VERIFY) 用「证据梯度」溯源到一手;量级类必须找到数字出处
4. 判定(LABEL) 给每条打 ✅ / 🟡 / 🔴 / ⚠️ 标签 + 写出一手源
5. 落档(LOG) 输出核验矩阵;投资场景同步纠正知识库 + 记 log(含纠错痕迹)
证据可信度梯度(高 → 低)
| 级 | 来源 | 用法 |
|---|---|---|
| S | SEC/EDGAR、SEDAR+、公司 IR / 官方 PR、8-K/10-Q/年报、标准文本(如 MSA 规范)、公司官网产品页、实时行情 API | 唯一可作"事实"的源 |
| A | 一线行业媒体、券商研报(署名)、原始论文 | 可作旁证,需标"分析师估/媒体口径" |
| B | 博客、substack、雪球/Stocktwits、转述新闻 | 仅作线索,必须回溯到 S/A |
| C | 匿名推文 / 小作文 / 群消息 / 其他 AI 的输出 | 只能当"待查线索清单",本身不是证据 |
"其他 AI 给的结论"= C 级。把它当成一份待核实的 leads,逐条降级查源,绝不能因为"AI 也这么说"就当佐证。
8 类常见失真模式(查到任一 → 标 🔴/⚠️)
- 张冠李戴 / 误归属:把 A 公司的订单/合作安到 B 头上。→ 查原始 PR 的主体到底是谁。
- 循环引用:媒体甲引媒体乙、乙又引当事人本人。→ 顺着引用链找最初出处,看是否独立。
- 推断升级成事实:"未具名客户"被脑补成具体公司;"潜在"被说成"已落地"。→ 区分披露事实 vs 社区推断。
- 选择性取利好/利空:只讲一个事实的有利面(如"被收购=利好",隐藏"订单被取消")。→ 同一事实强制双向看。
- 过时数据:拿 N 年前的 PR / 旧估值当现状。→ 检查日期;找最新一手覆盖。
- 营销展示 ≠ 商业事实(双向):官网 logo 墙 ≠ 供货合同;logo 撤了也 ≠ 关系归零。→ 用合同/财报判断,不靠营销页;缺失是弱证据,不过度解读("搜不到≠不存在")。
- 量级未证实:涨幅/估值/出货量/市值等数字没出处。→ 找财报/行情 API 实测;对不上就标 ⚠️。
- 概念混淆:把同名不同义、同公司不同产品线、相近标准混为一谈。→ 回到定义/规格书逐项对齐。
判定标签
| 标签 | 含义 | 必须附 |
|---|---|---|
| ✅ 证实 | S 级一手源确认 | 源链接/出处 + 关键数字 |
| 🟡 部分/需 nuance | 方向对但有偏差或前提 | 说明偏在哪 |
| 🔴 错误/误导 | 与一手源冲突 / 张冠李戴 | 正确事实 + 源 |
| ⚠️ 未证实 | 找不到一手源 | 注明"已查 X 未果,存疑" |
关系强度分级(判断"A 和 B 有没有关系"时)
合同/8-K > 入股 > 战略合作 PR > demo/样品/qual > 论坛传闻。
不能把弱级别说成强级别(如把"OFC demo"写成"量产收入")。
下"没有关系"结论前,先穷尽 vector 路径(见 agent-tool-escalation Case 6)。
下结论前的自检三问
- 我现在要采信/反驳的这条,源是 S/A 还是 B/C? C 级没回溯到 S/A → 只能标 ⚠️,不能下定论。
- 这是披露的事实,还是别人的推断/口径?量级类有没有数字出处?
- 我是不是因为"听起来合理 / 用户这么说 / 别的 AI 也这么说"就想认同?plausible ≠ correct——验了再说。
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 85 lines · 162 tokens per session scan A 1da81f2f6fe0
claim-verification is a skill published in the GitHub repository realnaka/alphaloop (19 stars, last pushed 2mo ago), licensed MIT. It adds 162 tokens to every session and 2,036 once invoked, about $0.0008 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
roblox-input
Use when handling Roblox keyboard, mouse, gamepad, touch, motion input, or cross-platform action binding.
roblox-lighting
Use for Roblox lighting, atmosphere, day/night, or post-processing effects.
roblox-luau-core
Use for Luau language semantics, tables, control flow, string patterns, scope, closures, and cross-language translation errors.
roblox-performance
Use when profiling Roblox performance or diagnosing FPS, memory, network, mobile, or hot-path problems.
roblox-server-data
Use for Roblox server or cross-server data: OrderedDataStore leaderboards, MessagingService, world state, seasons, or guilds.
changelog-detective
Detect what changed in the product by comparing before and after a deploy or date — find new behaviors, new errors, changed user flows, and unexpected side effects that weren't in the release notes. Use after a deploy or when the user suspects something changed but doesn't know what.