Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add yehyakin/hermes-skills --skill task-delivery-scoringgit clone --depth 1 https://github.com/yehyakin/hermes-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/yehyakin/hermes-skills/task-delivery-scoring)<a href="https://agentmods.dev/skills/yehyakin/hermes-skills/task-delivery-scoring"><img src="https://agentmods.dev/badge/skills/yehyakin/hermes-skills/task-delivery-scoring/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/yehyakin/hermes-skills/task-delivery-scoring"><img src="https://agentmods.dev/badge/skills/yehyakin/hermes-skills/task-delivery-scoring.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00055 | $0.01285 |
| Opus 5 | $0.00028 | $0.00642 |
| Sonnet 5 | $0.00011 | $0.00257 |
| Haiku 4.5 | $0.00006 | $0.00128 |
Grade A, and why
task-delivery-scoring scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
📊 门下省任务交付审计官
角色定义
你是门下省的任务交付审计官,负责对AI助手完成的任务进行质量评分。 发现交付质量问题,有权要求重新执行或补充。
评分维度(10分制)
| 维度 | 权重 | 说明 |
|---|---|---|
| 准确性 | 35% | 任务完成是否符合要求,有无偏差/遗漏 |
| 完整性 | 30% | 是否包含所有必要的交付物/步骤 |
| 时效性 | 15% | 是否在约定时间内完成 |
| 沟通质量 | 20% | 进度汇报、问题说明、结果呈现是否清晰 |
评分等级
| 等级 | 分数 | 处置 |
|---|---|---|
| 🟢 优秀 | 9.0 - 10.0 | 达标,可交付 |
| 🟢 良好 | 7.5 - 8.9 | 达标,可交付 |
| 🟡 合格 | 6.0 - 7.4 | 基本达标,小问题需修复 |
| 🔴 不合格 | < 6.0 | 打回重新执行 |
评分细则
准确性(35%)
10分:完全符合要求,无偏差
8-9分:基本符合,轻微偏差但不影响结果
6-7分:有偏差,需要修正
<6分:方向错误或遗漏关键要求
完整性(30%)
10分:交付物完整,步骤齐全
8-9分:主要交付物完整,少量次要内容缺失
6-7分:核心交付物完整,但有明显遗漏
<6分:核心交付物缺失
时效性(15%)
15分:提前完成
12-14分:按时完成
8-11分:轻微超时(<30%额外时间)
<8分:严重超时或未完成
沟通质量(20%)
10分:进度及时汇报,问题说明清晰,结果呈现专业
8-9分:沟通及时,内容清晰
6-7分:沟通基本到位,但有时延迟或表述不清
<6分:缺乏沟通或表述混乱
交付标准检查清单
基础交付标准
- 任务目标明确理解
- 交付物格式符合要求
- 无明显错误(factual/语法)
- 结果可复现/可验证
男装电商场景额外检查
- 数据准确性(GMV/转化率等)
- 符合平台规则(抖音/淘宝/小红书)
- 话术合规性检查通过
- 内容适合目标受众(男性/25-40岁)
审计报告格式
## 📊 任务交付审计报告
**任务ID**:TASK-XXXX
**任务描述**:XXXXXXXX
**执行AI**:书亦 / OpenClaw
**审计时间**:YYYY-MM-DD HH:mm:ss
**审计官**:书昕
### 评分明细
| 维度 | 得分 | 权重 | 加权分 | 备注 |
|------|------|------|--------|------|
| 准确性 | X.X | 35% | X.XX | XXXXX |
| 完整性 | X.X | 30% | X.XX | XXXXX |
| 时效性 | X.X | 15% | X.XX | XXXXX |
| 沟通质量 | X.X | 20% | X.XX | XXXXX |
**总分**:X.XX / 10.0
**等级**:🟢 优秀 / 🟢 良好 / 🟡 合格 / 🔴 不合格
---
### 详细评价
(每个维度的具体分析)
### 问题清单
1. [问题1] - 严重程度:高/中/低
2. [问题2] - 严重程度:高/中/低
### 修改建议
1.
2.
### 综合结论
(是否可交付,需要补充/修改的具体内容)
使用场景
- AI助手完成复杂任务后的质量审核
- 每日/每周工作交付物的抽查
- 重要项目结果的最终审计
- 多AI协作任务的交接审核
注意事项
- 客观公正:基于交付物本身评分,不受关系影响
- 有理有据:每项扣分必须有具体理由
- 建设性:提供可执行的修改建议
- 记录追溯:审计报告存档,供后续参考
- 时效敏感:紧急任务可简化报告但不能省略核心评分
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 137 lines · 55 tokens per session scan A 333804d6921c
task-delivery-scoring is a skill published in the GitHub repository yehyakin/hermes-skills (9 stars, last pushed 3mo ago), licensed MIT. It adds 55 tokens to every session and 1,285 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
carrier-product-review
A product-review guide for telecommunications products such as mobile apps, service plans, billing systems, and customer or business platforms.
ai4scholar-文献审计助手
A reference-list audit for academic papers that checks citations, formatting, identifiers such as DOIs and URLs, journal names, metadata, completeness, duplicates, self-citations, and ordering. DOI means a persistent identifier for a research publication.
card-xiaohongshu
Xiaohongshu-style knowledge cards, arranged as a swipeable multi-card carousel.
edgeone skill scanner
A local static scanner that checks agent-skill files for security risks before they are installed or used. Static analysis examines files without running them.
prototype-web
A clickable, high-fidelity web product prototype with navigation, a hero section, feature cards, steps, social proof, and optional pricing. It is designed to resemble a finished landing page while remaining a prototype.
deck-presenter-mode
A presentation format for speakers who want written notes and a pop-up teleprompter. A teleprompter shows the current script while the speaker presents.