Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/xu-xiang/everything-claude-code-zh/evalgit clone --depth 1 https://github.com/xu-xiang/everything-claude-code-zhWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/xu-xiang/everything-claude-code-zh/eval)<a href="https://agentmods.dev/commands/xu-xiang/everything-claude-code-zh/eval"><img src="https://agentmods.dev/badge/commands/xu-xiang/everything-claude-code-zh/eval.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00005 | $0.00565 |
| Opus 5 | $0.00003 | $0.00282 |
| Sonnet 5 | $0.00001 | $0.00113 |
| Haiku 4.5 | $0.00001 | $0.00056 |
Grade A, and why
eval scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Eval 命令
根据准入标准(Acceptance Criteria)评估实现情况:$ARGUMENTS
你的任务
执行结构化评测(Evaluation),验证实现是否符合需求。
评测框架(Evaluation Framework)
评分器(Grader)类型
-
二元评分器(Binary Grader) - 通过/失败(Pass/Fail)
- 是否工作?是/否
- 适用场景:功能完成度、缺陷修复(Bug Fixes)
-
标量评分器(Scalar Grader) - 分数 0-100
- 工作效果如何?
- 适用场景:性能、质量指标
-
量规评分器(Rubric Grader) - 维度评分
- 多维度评估
- 适用场景:全面审查
评测流程
步骤 1:定义标准
Acceptance Criteria:
1. [标准 1] - [权重]
2. [标准 2] - [权重]
3. [标准 3] - [权重]
步骤 2:运行测试
针对每个标准:
- 执行相关测试
- 收集证据
- 评分结果
步骤 3:计算得分
Final Score = Σ (criterion_score × weight) / total_weight
步骤 4:生成报告
评测报告
总体情况:[通过/失败] (得分: X/100)
标准细分
| 标准 (Criterion) | 得分 | 权重 | 加权分 |
|---|---|---|---|
| [标准 1] | X/10 | 30% | X |
| [标准 2] | X/10 | 40% | X |
| [标准 3] | X/10 | 30% | X |
证据(Evidence)
标准 1:[名称]
- 测试:[测试内容]
- 结果:[产出结果]
- 证据:[截图、日志、输出]
建议
[如果未通过,需要改进的地方]
Pass@K 指标
针对非确定性评测:
- 运行 K 次
- 计算通过率
- 报告:"Pass@K = X/K"
提示:在标记功能完成之前,建议使用 eval 进行准入测试(Acceptance Testing)。
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 89 lines · 5 tokens per session scan A 3a0dcb249746
eval is a command published in the GitHub repository xu-xiang/everything-claude-code-zh (1,927 stars, last pushed 6mo ago), licensed MIT. It adds 5 tokens to every session and 565 once invoked, about $0.0000 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
awesome-chatgpt
Search awesome-ChatGPT-repositories for open-source GitHub repositories related to ChatGPT and LLMs.
init
Scaffold a new MindBase project (v2 layout). Usage: /mb:init [template] [-- mission ...].
commit
智能生成 Git 提交信息并提交.
pr
Handle the full workflow from current branch state to an open, CI-monitored pull request.
doctor.es
Diagnostica problemas de inferencia LLM en Mac: asiai doctor verifica el estado de los motores, conflictos de puertos, carga de modelos y estado de la GPU.
requirement-review
需求文档多角色评审(requirement-review):需求文档 → 7-Agent 并行评审 → 重构高质量需求文档(Runtime 受控流程,0-7 阶段状态机).