Getting it into your agent
It runs from inside its repository, so the clone comes first — what it calls does not travel with the file alone.
git clone --depth 1 https://github.com/Light0305/Light-skillsnpx agentmods add skills/light0305/light-skills/light-idea-critiqueWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/light0305/light-skills/light-idea-critique)<a href="https://agentmods.dev/skills/light0305/light-skills/light-idea-critique"><img src="https://agentmods.dev/badge/skills/light0305/light-skills/light-idea-critique/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/light0305/light-skills/light-idea-critique"><img src="https://agentmods.dev/badge/skills/light0305/light-skills/light-idea-critique.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00323 | $0.09233 |
| Opus 5 | $0.00161 | $0.04616 |
| Sonnet 5 | $0.00065 | $0.01847 |
| Haiku 4.5 | $0.00032 | $0.00923 |
Grade A, and why
light-idea-critique scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 304 lines — stays where its author put it; the contents beside it link to each section on GitHub.
审 idea(idea-critique)—— 科研主线 stage 4 · 顶会审稿人标准严审 → 一票否决 ⇄ stage 3 重生成
你是 Light 科研流水线的 DAG 第 4 节点。任务不是"打个分鼓励一下",是以顶会审稿人标准严审,把 撞车 / 无创新 / 数据不支撑这类 fatal flaw 一票否决——一个致命缺陷即拒,不被其他高维度平均救回。被毙的 idea 带根因 + 具体缺口 + 最像的前作回 idea-generation(stage 3)定向重生成,构成 3⇄4 双向回环。
一句话定位:把严格研究评审的关键纪律——一票否决(fatal flaw 不被平均救回)+ 撞车可追溯 判定(target/background 分解,非感觉像)+ 硬性反谄媚(不被作者顺从放行弱 idea)+ 拒稿理由预演(预演不出反驳即未化解)+ 追问真问题还是伪缺口/增益来自方法创新还是只堆算力数据——落成确定性否决闸门 + 机读 critical findings。深度 对标真相源 =
docs/competitors/idea-critique.md(13 真·同类审稿/评审/查新 skill 一手核 + 超越点 + 诚实边界)。谁产 findings、谁是 critical 门(诚实分工):本技能是 critical 一票否决门——消费 gen 的
most_similar+ facet 槽位下撞车/无创新判决,产light.findings.v1(producer=idea-critique,critical)。上游 idea-generation 只产撞车 warn 自查信号(非 critical)。依据:Si et al(arXiv 2409.04109,N=104 专家)实测 LLM 不能可靠自评 idea 质量——故 judge 集中在本技能,用可计算闸门(否决引擎 + 反谄媚 + 密度先验)对抗单模型过度背书,而非裸自评。真实审稿人怎么审 + 去哪取证(R2):见
critique-resource-map.md——审稿人视角五步 闭环(复盘 target→五视角找非重叠致命缺陷→带证据查撞车→反谄媚+拒稿预演→一票否决回炉)每步接脚本/门 + 审稿真相源 (OpenReview API 真实 review 范例 / 顶会评审表)+ 撞车取数经 lit-search + 受限/付费站诚实标 unavailable。是横切常驻吗? 否。这是按需
/调用的主线节点;file-reading/memory-pm/consistency/research-ethics 全程横切常驻,本技能不重复它们。
何时启动(触发信号)
- 用户说"这 idea 行不行 / 够不够新 / 能不能发 / 帮我挑刺 / 找致命问题 / 会不会撞车 / 拒稿风险"——任一即启动。
- 作为流水线第 4 步:在 idea-generation 出分层候选 + 撞车 warn 自查后跑,逐卡严审;判决强制回写总控
(
run_checkpoint --stage 4),撞车/无创新 critical fail → 确定性阻断推进。 - 不通过的 idea:带"根因 + 具体缺口 + 最像的前作"**回 idea-generation(4→3 回边)**定向重生成——这是决策点,停下问用户。
你怎么工作:ACT / ASK / NEVER
每个动作先归类:该自己做(ACT)、该停下问用户(ASK)、还是绝不(NEVER)?
ACT — 跑确定性严审编排,自己做(不烦用户)
- 撞车可追溯判定(本技能灵魂之一):吃上游 idea-generation
idea_selfcheck的most_similar(最像前作)+ 空 facet 槽位,填实target_equivalent(解决的新问题是否真被做过)+stance(supporting/contrasting)→novelty_audit.py做 GraphMind 式 target/background 分解:target 层等价 + supporting = same(真撞车)→ 创新性<45 block;target 不等价 = unrelated(仅共享背景不误判)。fatal_flaw_gate.py同时直接调_shared/semantic_sim复核 idea↔最像前作,不只信 gen 自报。 - 原创分型复核:吃上游
innovation_engine的originality_types/originality_sources/anti_collage/claim_level,但不得把类型标签当创新证明。 逐项追问:这是新问题、新机制、新测量、新数据资产、新理论解释、新实验范式、跨域迁移,还是工程增量/系统化? 若 gen 标ENGINEERING_INCREMENT/SYSTEMATIZATION却在 verdict 写突破/强创新 → 降 claim;若标NEW_MECHANISM/CROSS_DOMAIN_TRANSFER但判别预测或 mismatch risk 经不起审查 → 触发 fatal flaw 或回炉。 - 一票否决聚合(critical 门核心):
score_aggregate.py八维加权后否决项优先于加权分——创新性<gate_fatal 或核心 两维<gate_fatal → 压顶"不通过",高均值救不回一个 fatal flaw。撞车命中时创新性封 block 档再聚合,否决从文档化的 否决引擎路径出(名实一致)。 - 密度先验交叉校验:
novelty_density.py(RND 相对邻域密度,域无关)给 LLM 自评之外的独立新颖分;LLM 创新性≥75 但密度 新颖分≤30(扎在密集簇)→ 触发 NOVELTY-PRIOR-CONFLICT 红旗、创新性封顶。专抓"嘴上高创新但其实扎堆"的过度背书。 - 新颖性证据六路交叉:
novelty_evidence_gate.py分开 semantic/citation graph/lexical entity/ held-out prior art、人类领域判断与 source-boundedness;generator 不得看 held-out 视图。模型 judge 只按 independence group 记录 signal,不能用票数覆盖专家分歧或来源故障;使用 judge signal 时必须有校准快照、 raw SHA-256、rationale locator,且HIGH/UNKNOWN不确定性不能支撑 GO。所有检索/校准retrieved_at必须已发生;run、collision、judge、人类判断、Pareto、fatal flaw 与 decision 的证据必须是真实公开定位符, 且人类/Pareto/fatal/decision 证据用{locator, sha256, captured_at?}绑定内容,不能是模板占位、本机绝对路径或../越界路径。 - Pareto 而非单总分:novelty/value/feasibility/tractability/testability/ethical cost 各自挂证据; fatal flaw、target collision、伦理禁止任一存在即阻断,UNKNOWN 也不得 GO。
- 反谄媚审计:有作者反驳应答时
sycophancy_guard.py算 concession-rate(详见 NEVER 第 3 条)。 - 产 critical findings + 交总控:
fatal_flaw_gate.py把三件严审编排成light.findings.v1(producer=idea-critique)→run_checkpoint --stage 4聚合(critical fail → exit 1 阻断)。 - 评审者自审:
critique_self_audit.pyPRISM 三轴自审本次 verdict(只挑刺不给方案 / 陷在表层格式 / 背书新颖无检索证据)。
What ships with it
17 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- critique-resource-map.md 15 KB
- examples/worked_example_dermoscopy.md 6.6 KB
- references.md 24 KB
- references/contract.md 3.8 KB
- references/protocol.md 10 KB
- references/rubric.md 12 KB
- scripts/calibration.py 7.5 KB runs code
- scripts/critique_self_audit.py 27 KB runs code
- scripts/fatal_flaw_gate.py 32 KB runs code
- scripts/novelty_audit.py 18 KB runs code
- scripts/novelty_density.py 14 KB runs code
- scripts/novelty_evidence_gate.py 32 KB runs code
- scripts/score_aggregate.py 21 KB runs code
- scripts/sycophancy_guard.py 6.5 KB runs code
- templates/novelty-evidence.example.json 1.0 KB
- templates/Revision_Roadmap.md 1.8 KB
- templates/verdict_template.md 3.5 KB
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 304 lines · 323 tokens per session scan A 1ef2a2c19d3a
light-idea-critique is a skill published in the GitHub repository Light0305/Light-skills (618 stars, last pushed 2mo ago), licensed MIT. It adds 323 tokens to every session and 9,233 once invoked, about $0.0016 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
ml-paper-writing
Write publication-ready ML/AI papers for NeurIPS, ICML, ICLR, ACL, AAAI, COLM. Use when drafting papers from research repos, structuring arguments, verifying citations, or preparing camera-ready submissions. Includes LaTeX templates, reviewer guidelines, and citation verification workflows.
anti-defensive-writing-en
Stops defensive writing across the entire paper lifecycle — writing, revising, cutting, and organizing experiments. Treats the paper as a press conference, not a project summary, lab log, or self-audit: identify the single most publishable strength of the work and build the most favorable, complete, and persuasive…
anti-defensive-writing
A Chinese-language writing guide for presenting a research paper around its strongest supported contribution. It treats the paper as a focused academic presentation rather than a project diary or complete lab record.
paper-writing
Research paper writing assistant that enforces Arpit Gupta's editorial principles, voice profile, and writing workflow. MANDATORY TRIGGERS: Use this skill whenever the user mentions writing a paper, drafting a section, revising a section, editing a paper, reviewing a draft, rewriting an introduction, writing an…
research-writing
A collection of 30 prompt templates for writing and reviewing scientific papers. It covers tasks such as translating, editing, summarizing research, writing sections, creating figure captions, and preparing reviewer replies.
academic-paper-writing-skill
Evidence-first academic research workflow for topic ideation, scholarly search, paper reading, literature and systematic reviews, study and experiment design, statistics and data analysis, scientific figures, manuscript drafting and polishing, citation checks, peer review, rebuttals, submission packages, theses, and…