gl-thesis-variable-selection

A Chinese-language research-planning skill for checking AI-generated variable choices in undergraduate and master’s business or economics theses. It turns concepts in an approved topic into measurable variables with proposed formulas, data sources, literature checks, and control-variable choices.

In plain words
What is it for?
Use it to review a variable plan, identify outcome and explanatory variables, choose main and alternative measures, assess data feasibility, select controls from empirical papers, and produce a variable table or control-variable matrix.
Why use it?
It helps catch differences between similarly named measures, unsuitable controls, data-availability problems, and unsupported claims. It does not replace real literature checks or run the final regression.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/ali66611/glskill/gl-thesis-variable-selection
Any agent
npx skills add Ali66611/glskill --skill gl-thesis-variable-selection
Clone the repo
git clone --depth 1 https://github.com/Ali66611/glskill

Made for: Claude Code, Codex.

Per session 129 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,598 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00129 $0.02598
Opus 5 $0.00064 $0.01299
Sonnet 5 $0.00026 $0.00520
Haiku 4.5 $0.00013 $0.00260

Measured 2d ago against content hash 544a2b8a71ed, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

gl-thesis-variable-selection scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

gl-thesis-variable-selection/SKILL.md · 152 lines

How it starts

The opening of the file, as written. The whole thing — 152 lines — stays where its author put it; the contents beside it link to each section on GitHub.

GL 论文选题变量确定 Skill

核心任务

把“题目里的概念”转成有文献依据、有数据来源、能复现的变量方案。

核心顺序:

先分变量角色
→ 再查衡量口径
→ 再过数据闸门
→ 再筛控制变量
→ 最后形成研究方案

使用前提醒

本 Skill 需要检索真实文献来核验变量口径和控制变量,建议本机提前安装并启用:

  • CNKI Skill 套件:用于检索中文经管文献、硕博论文和变量衡量方式,例如 cnki-searchcnki-paper-detail
  • GS Skill 套件:用于检索 Google Scholar 英文文献、机制和替代口径,例如 gs-searchgs-fulltext
  • 文档读取工具:用于读取用户提供的 PDF、Word 和 Excel 文献。

开始任务时先检查上述工具是否可用。若 CNKI Skill 或 GS Skill 未安装,必须明确提醒使用者先安装相关 Skill;在工具恢复前只能生成初步候选方案,所有未经真实文献核验的结论必须标注“待核验”,不得伪造文献依据。

输入检查

先检查 CNKI Skill、GS Skill 和文档读取工具是否可用,再读取用户提供的题目、学历、专业、论文用途、10 篇相关实证文献和可用数据源。

信息不足时最多问 5 个关键问题。若用户未提供 10 篇文献且当前环境有 CNKI、Google Scholar 或文献检索工具,主动补充检索;工具不可用时标注“待核验”,不得编造。

执行流程

  1. 审查已有 AI 答案:用户提供普通 AI 的变量方案时,先读取 AI 变量答案审查规则,检查证据、角色、口径、固定效应、内生性和数据落地风险。
  2. 分变量角色:识别 Y、X、控制变量、机制变量、调节变量和异质性分组。机制变量或后处理变量不得机械放入基准控制变量。
  3. 确定核心变量口径:读取 核心变量衡量规则,输出主衡量方式、替代衡量方式、公式、文献和数据来源。
  4. 评估数据可得性:读取 数据可得性评分,完成六维评分和红黄绿闸门。
  5. 筛选控制变量:读取 控制变量筛选规则,从 10 篇文献抽取、对齐、计频并检查理论必要性。
  6. 生成交付物:读取 输出契约,生成变量设计报告和控制变量矩阵。需要 Excel 时复制 assets/控制变量文献矩阵模板.xlsx,不修改原始模板。

硬规则

  • 变量名称相似不等于衡量口径相同;ROA/ROE、企业年龄/上市年限、Top1/Top10 不得强行合并。
  • 高频控制变量只代表文献共识信号,不代表自动必选。
  • size、lev 是企业层面高频基础候选,不是所有研究的绝对必选。
  • 8-10 个控制变量是常见起点,不是数量指标;不得为了凑数加入无理论依据的变量。
  • 固定效应不计入控制变量数量。
  • 企业面板首先检查企业固定效应与年份固定效应;加入企业固定效应后,不机械叠加基本不变的行业和地区固定效应。
  • 滞后一期不能自动解决内生性;工具变量必须单独论证相关性、排除性和数据可行性。
  • 硕士变量可以多源合并或构建综合指数,但不是越复杂越好。
  • 熵值法权重表示样本信息差异,不等于理论重要性。
  • 爬虫数据并非天然不被认可,但技术、合规、清洗和复现风险通常更高。
  • 不为追求显著性先定结论,再倒推变量或控制变量。
  • 任何未从真实文献或数据源确认的内容都标注“待核验”。

工具路由

  • 中文变量口径与控制变量文献:优先 CNKI、中文权威经管期刊、CSSCI 和北大核心。
  • 英文前沿、机制与替代口径:使用 Google Scholar 或英文文献检索工具补充。
  • 表格交付:优先生成 .xlsx;同时在 Markdown 报告中给出可预览的摘要表。
  • 用户提供 PDF、Word 或 Excel 时,先读取原文件,不凭标题猜测变量定义。

中文核心文献检索与全文核验流程

中文核心论文不能只凭搜索结果摘要或题名进入控制变量矩阵。按以下顺序执行,并把每一步的证据留存到项目文件中。

1. 先筛期刊层级,再筛题目相关性

  1. 在 CNKI 高级检索中使用题名或主题组合检索,例如“数字化转型 AND 绿色创新”。
  2. 在来源限定中选择“CSSCI”和“北大核心”相关来源;如果页面把两者作为并集筛选,报告中写明“至少命中其中一种”,不要笼统写成“顶刊”。
  3. 优先保留同一被解释变量、同一数据层级和相近样本的中文核心实证论文;可按被引量排序用于发现高影响文献,但被引量不是核心级别证明。
  4. 每篇候选必须核对作者、题名、期刊、年份、卷期、页码和 DOI。题名相同的论文必须继续核对作者和期刊,不得按题名唯一匹配。

Read the full file on GitHub · 152 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 152 lines · 129 tokens per session scan A 544a2b8a71ed

Subscribe to this mod's changes

gl-thesis-variable-selection is a skill published in the GitHub repository Ali66611/glskill (6 stars, last pushed 1mo ago), licensed MIT. It adds 129 tokens to every session and 2,598 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.