Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add serejaris/kimi-skills --skill auto-stat-testgit clone --depth 1 https://github.com/serejaris/kimi-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/serejaris/kimi-skills/auto-stat-test)<a href="https://agentmods.dev/skills/serejaris/kimi-skills/auto-stat-test"><img src="https://agentmods.dev/badge/skills/serejaris/kimi-skills/auto-stat-test/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/serejaris/kimi-skills/auto-stat-test"><img src="https://agentmods.dev/badge/skills/serejaris/kimi-skills/auto-stat-test.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00121 | $0.01546 |
| Opus 5 | $0.00060 | $0.00773 |
| Sonnet 5 | $0.00024 | $0.00309 |
| Haiku 4.5 | $0.00012 | $0.00155 |
Grade A, and why
auto-stat-test scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 128 lines — stays where its author put it; the contents beside it link to each section on GitHub.
statistical-test-suite
自动统计检验工具 —— 根据数据特征自动选择合适的统计检验方法(t 检验 / 卡方 / ANOVA / Mann-Whitney 等),一键输出检验结果和中文通俗解读。
能力概览
| 功能 | 说明 |
|---|---|
| 独立样本 t 检验 | 2 组 + 正态数据,比较均值差异 |
| Welch t 检验 | 2 组 + 正态但方差不齐 |
| Mann-Whitney U | 2 组 + 非正态数据(非参数) |
| 单因素 ANOVA | 3+ 组 + 正态数据 |
| Kruskal-Wallis | 3+ 组 + 非正态数据(非参数) |
| 卡方独立性检验 | 两个分类变量的关联性 |
| 配对 t 检验 | 前后对比(正态) |
| Wilcoxon 符号秩 | 前后对比(非参数) |
| 自动选择 | 根据组数、正态性、数据类型自动决定 |
| 通俗解读 | 每个指标和结论都给出中文白话说明 |
Quick Start
# 分组比较(自动选择检验方法)
python3 scripts/statistical_test_suite.py data.csv --group treatment --value score
# 卡方检验(两个分类变量)
python3 scripts/statistical_test_suite.py survey.csv --group gender --value preference
# 配对检验(前后对比)
python3 scripts/statistical_test_suite.py experiment.csv --col1 pre_score --col2 post_score --paired
# 强制指定检验方法
python3 scripts/statistical_test_suite.py data.csv --group group --value score --test mann-whitney
# 保存结果到 JSON
python3 scripts/statistical_test_suite.py data.csv -g treatment -v score -o result.json
详细用法
模式一:分组比较
用 --group 指定分组列,--value 指定比较列,工具自动判断用哪种检验。
python3 scripts/statistical_test_suite.py <数据文件> --group <分组列> --value <数值列> [选项]
自动选择逻辑:
- 两列都是分类变量 → 卡方检验
- 2 个组 + 数据正态 → 独立样本 t 检验(方差不齐则用 Welch t)
- 2 个组 + 数据非正态 → Mann-Whitney U 检验
- 3+ 个组 + 数据正态 → 单因素 ANOVA
- 3+ 个组 + 数据非正态 → Kruskal-Wallis 检验
模式二:配对比较
用 --col1 和 --col2 指定前后两个变量列。
python3 scripts/statistical_test_suite.py <数据文件> --col1 <前> --col2 <后> --paired [选项]
自动选择逻辑:
- 差值正态 → 配对 t 检验
- 差值非正态 → Wilcoxon 符号秩检验
参数说明
| 参数 | 缩写 | 必填 | 默认值 | 说明 |
|---|---|---|---|---|
input |
— | 是 | — | 输入文件路径(CSV/TSV/Excel/JSON) |
--group |
-g |
模式一 | — | 分组变量列名 |
--value |
-v |
模式一 | — | 数值/分类变量列名 |
--col1 |
— | 模式二 | — | 配对检验第 1 个变量列名 |
--col2 |
— | 模式二 | — | 配对检验第 2 个变量列名 |
--paired |
— | 否 | false |
启用配对检验模式 |
--test |
-T |
否 | 自动 | 强制检验方法(见下方列表) |
--alpha |
-a |
否 | 0.05 |
显著性水平 |
--output |
-o |
否 | 标准输出 | 结果 JSON 保存路径 |
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 128 lines · 121 tokens per session scan A 4e32c985d755
auto-stat-test is a skill published in the GitHub repository serejaris/kimi-skills (6 stars, last pushed 1mo ago), licensed MIT. It adds 121 tokens to every session and 1,546 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
literature-searcher
Search CrossRef, OpenAlex, PubMed, Semantic Scholar, and optional Scopus; deduplicate results, download open-access PDFs by DOI, classify papers, monitor new results, and analyze coverage. Use when asked to search literature, monitor a topic, download an open-access paper, classify papers, or analyze literature gaps.
科研可视化工具
A research-visualisation workflow for examining data and producing publication-ready charts for scientific papers.
math-modeling
A workflow for mathematical modelling, where real-world questions are represented with mathematics and solved with code or analysis.
math-modeling
A workflow for mathematical modelling, where real-world questions are represented with mathematics and solved with code or analysis.
LaTeX工具
A tool for creating, compiling, and checking mathematical modelling papers written in LaTeX, a document system often used for technical writing.
aql-authoring
This skill should be used when the user asks to "write an AQL query", "optimize an AQL query", "review AQL", or "query openEHR data" — the multi-step authoring/optimization workflow for AQL (Archetype Query Language) over openEHR clinical data. For a one-off explanation of an existing query or a single AQL…