Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add BiglionX/SkillHub --skill skill-smoke-testgit clone --depth 1 https://github.com/BiglionX/SkillHubWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/biglionx/skillhub/skill-smoke-test)<a href="https://agentmods.dev/skills/biglionx/skillhub/skill-smoke-test"><img src="https://agentmods.dev/badge/skills/biglionx/skillhub/skill-smoke-test/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/biglionx/skillhub/skill-smoke-test"><img src="https://agentmods.dev/badge/skills/biglionx/skillhub/skill-smoke-test.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00068 | $0.00689 |
| Opus 5 | $0.00034 | $0.00345 |
| Sonnet 5 | $0.00014 | $0.00138 |
| Haiku 4.5 | $0.00007 | $0.00069 |
Grade A, and why
skill-smoke-test scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
技能包冒烟测试 (skill-smoke-test)
本技能指导 Agent 在隔离沙箱中对技能包做运行时冒烟测试。目标:确认 SKILL.md 能被 Agent 正确加载,且技能在代表性任务上产生预期输出。
何时使用
- 发布前运行时验证(配合 skill-package-validator 的静态校验)
- 审核第三方/爬取来的技能包(GitHub 全球搜索入库前)
- 复现用户报告的"技能不工作"问题
工作流程
1. 准备沙箱工作区
- 创建临时目录(如
.tmp-smoke/<skill-name>-<ts>),不得污染仓库 - 将技能包完整复制进去(SKILL.md + scripts/ + assets/)
- 确认沙箱无网络敏感操作(如需联网,记录并限制超时)
2. 加载技能
- 把 SKILL.md 作为技能目录加载进 Agent harness(DSH / DeerFlow / Claude Code 均可)
- 确认 frontmatter 解析成功(name/description 正确呈现)
3. 设计代表性任务(1-3 个)
- 从 description 与正文提取技能的典型使用场景
- 任务要可判定:有明确期望输出(文件生成/答案正确/格式合规)
- 至少 1 个边界/异常任务(缺输入、非法参数)
4. 执行
- 逐个任务运行,记录:命令/提示、超时、退出码、输出摘要
- 注意资源占用(CPU/内存/磁盘)与权限越界(禁止写仓库外路径、禁止删除)
5. 检查与报告
- 对照期望输出判定 PASS/FAIL
- 生成报告(markdown):
# 冒烟测试报告: <skill-name>@<version>
- 环境: <harness+模型>
- 任务1: PASS/FAIL (期望 vs 实际)
- 任务2: ...
- 异常/安全观察: ...
- 结论: 可用 / 需修复 (建议)
安全底线
- 只在沙箱目录内写文件;技能脚本若尝试越界 → 立即终止并标记"危险"
- 长任务设超时;循环/无限等待直接 kill
- 不把技能的真实凭据/密钥带入测试环境
参考
- 模式参考:
deer-flow/.agent/skills/smoke-test/(vendored DeerFlow 的冒烟技能,含报告模板) - 技能规范:Agent Skills 标准(SKILL.md frontmatter + 正文)
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 59 lines · 68 tokens per session scan A c550e9821550
skill-smoke-test is a skill published in the GitHub repository BiglionX/SkillHub (34 stars, last pushed 3d ago), licensed Apache-2.0. It adds 68 tokens to every session and 689 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
research-engineer
An uncompromising Academic Research Engineer. Operates with absolute scientific rigor, objective criticism, and zero flair. Focuses on theoretical correctness, formal verification, and optimal implementation across any required technology.
tika-eval-compare
Compare extracts from two Tika builds over a corpus to detect regressions in content, encoding, exceptions, and embedded-document handling. Use for "compare before/after extracts", "eval this change against the corpus".
neuron-evaluation-engineer
Create and run AI evaluations with datasets, assertions, and output drivers in Neuron AI. Use this skill whenever the user mentions evaluation, testing AI systems, creating evaluators, dataset-driven testing, assertion-based validation, or wants to measure AI system performance. Also trigger for tasks involving…
jetson-validate-image
Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.
atmos-validation
Validate Atmos projects, components, arbitrary JSON Schema inputs, EditorConfig, and GitHub Actions; use affected-file selection and native CI annotations.
skill-benchmark
Benchmark AI skill effectiveness by measuring implementation quality against legacy constraints.