Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add csmar432/finai-research --skill fin-review-loopgit clone --depth 1 https://github.com/csmar432/finai-researchWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/csmar432/finai-research/fin-review-loop)<a href="https://agentmods.dev/skills/csmar432/finai-research/fin-review-loop"><img src="https://agentmods.dev/badge/skills/csmar432/finai-research/fin-review-loop/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/csmar432/finai-research/fin-review-loop"><img src="https://agentmods.dev/badge/skills/csmar432/finai-research/fin-review-loop.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00071 | $0.02047 |
| Opus 5 | $0.00036 | $0.01024 |
| Sonnet 5 | $0.00014 | $0.00409 |
| Haiku 4.5 | $0.00007 | $0.00205 |
Grade A, and why
fin-review-loop scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 219 lines — stays where its author put it; the contents beside it link to each section on GitHub.
fin-review-loop
经济金融论文的对抗性review循环。对草稿进行多轮严格评审,检查实证严谨性、方法正确性、理论贡献和写作质量,给出可操作的修改建议。(AI review 不能替代同行评审,草稿必须经研究者核实后投稿。)
触发条件
- 关键词:
review评审审稿检查论文对抗性review论文检查 - Skill语法:
Skill: fin-review-loop
评分维度与权重
| 维度 | 权重 | 通过阈值 |
|---|---|---|
| 新颖性 (Novelty) | 30% | >= 6.0 |
| 实证严谨性 (Empirical Rigour) | 30% | >= 6.0 |
| 文献覆盖 (Literature Coverage) | 15% | >= 5.0 |
| 写作清晰 (Writing Clarity) | 15% | >= 5.0 |
| 学术影响 (Academic Impact) | 10% | >= 5.0 |
其中"写作清晰"维度必须包含 AI 味检测:全文不得出现 AI 典型句式 (详见 docs/writing-guide/ANTI_AI_WRITING_GUIDE.md), 结论段必须包含底气要素(具体数字/经济规模/机制描述/对比发现之一)。
评审难度级别
standard: 模拟标准学术审稿人strict: 模拟顶刊审稿人 (JF/JFE 级别)nightmare: 模拟严苛批评型审稿人 (如被拒稿后的防御性检查)
评审难度示例
standard
- 发现问题时会给出温和建议
- 接受主流方法选择
- 关注核心贡献是否清晰
strict
- 要求所有实证细节完备
- 质疑识别策略的每一步
- 检查文献是否覆盖最新顶刊
nightmare
- 预设论文会被拒,准备攻击
- 寻找方法论上的致命缺陷
- 模拟最严格的匿名审稿人
停止条件 (立即终止评审并报告用户)
满足以下任一条件时,立即停止评审:
- 新颖性 < 6.0 → 建议重新评估研究定位
- 实证严谨性 < 6.0 → 必须修复实证问题才能继续
- 已达到最大评审轮次 (4轮)
评审流程
第一步:解析论文
- 读取
output/fin-manuscript/下的所有.tex文件 - 提取论文结构:Introduction, Literature Review, Data, Methodology, Results, Conclusion
- 如文件不存在,扫描项目根目录和
papers/目录
第二步:诊断性检查
自动运行以下检查:
□ 平行趋势检验结果是否存在
□ 稳健性检验 >= 6 种
□ 异质性分析是否包含
□ 机制分析是否包含
□ 参考文献是否包含近3年顶刊论文
□ 变量定义表是否完整
□ 数据来源是否标注
□ 实证方法选择是否合理
第三步:逐维度评分
对每个维度进行 1-10 分评分,并说明理由:
| 维度 | 评分 | 理由 |
|---|---|---|
| 新颖性 | X | 边际贡献是什么?与现有文献区别? |
| 实证严谨性 | X | 识别策略是否合理?数据是否可靠? |
| 文献覆盖 | X | 是否覆盖最新顶刊?经典文献? |
| 写作清晰 | X | 逻辑是否清晰?论证是否连贯? |
| 学术影响 | X | 对该领域的潜在影响?引用潜力? |
第四步:生成逐节反馈
为论文每个章节生成具体、可操作的反馈:
### Introduction
- 问题: 边际贡献描述不够具体
- 建议: 明确说明与X论文的区别,本文的增量贡献是什么
### Data & Methodology
- 问题: 平行趋势图缺少统计显著性标注
- 建议: 在图中标注pre-treatment各期系数的置信区间
### Results
- 问题: 基准回归系数解读不够严谨
- 建议: 添加经济显著性解释(1个标准差变动对应Y变化X%)
第五步:识别审稿人攻击点
识别论文中最可能被审稿人攻击的弱点:
## 审稿人攻击点
1. [高风险] 审稿人会质疑平行趋势假设——需要pre-trends test p值
2. [中风险] 样本期间选择——为何选择2012-2022年?
3. [低风险] 稳健性检验中未包含安慰剂检验
第六步:生成修订计划
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 219 lines · 71 tokens per session scan A 88907741d0ae
fin-review-loop is a skill published in the GitHub repository csmar432/finai-research (100 stars, last pushed 2d ago), licensed MIT. It adds 71 tokens to every session and 2,047 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
r-econometrics
Run IV, DiD, and RDD analyses in R with proper diagnostics.
econometrics-phd-level
A guide to econometrics, the use of statistics to study relationships in data, based on a 12-part Korean lecture series. It routes questions to explanations of topics such as regression, panel data, instrumental variables, and causal comparisons.
did-causal
Use this Skill when the user needs to estimate causal treatment effects using difference-in-differences (DID) designs: two-way fixed effects (TWFE) regression, parallel trends pre-testing, Callaway-Sant'Anna staggered adoption estimator, and Goodman-Bacon decomposition. Covers both Python (linearmodels) and R (did…
r-econometrics
Generates rigorous, modern, reproducible R code for causal inference and panel econometrics with fixest, heterogeneity-robust DiD estimators (Callaway-Sant'Anna, Sun-Abraham, BJS, de Chaisemartin-D'Haultfoeuille), weak-IV-robust inference, optimal-bandwidth RDD via rdrobust, and wild cluster bootstrap. Use when the…
kami-deck
A lab-meeting deck on gut-microbiome links to sleep quality — the design, the results, the caveats, and the next experiment. Built as a decision-grade academic research deck for lab group, PI.
hps-academic-paper
A review deck on compositional generalization in large language models — the field map, the gap, the evidence, and open questions. Built as a decision-grade academic research deck for PI, lab group, reviewers.