Clowder AI is a self-hosted workspace where AI agents from different model families work together as a persistent team, retaining identities, shared evidence, and memory across tasks. It is for people who want to coordinate multiple AI agents without repeatedly rebuilding their context.
Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/zts212653/clowder-ai/eval-designnpx skills add zts212653/clowder-ai --skill eval-designgit clone --depth 1 https://github.com/zts212653/clowder-aiWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/zts212653/clowder-ai/eval-design)<a href="https://agentmods.dev/skills/zts212653/clowder-ai/eval-design"><img src="https://agentmods.dev/badge/skills/zts212653/clowder-ai/eval-design.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00090 | $0.04715 |
| Opus 5 | $0.00045 | $0.02357 |
| Sonnet 5 | $0.00018 | $0.00943 |
| Haiku 4.5 | $0.00009 | $0.00471 |
Grade A, and why
eval-design scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 211 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Eval Design — E0 资格门 · 出生证契约 · 运行拓扑 · 设计自检 · 五病体检 · 干预证
版本定位(2026-08-23 v2.2 修正轮):v0.3 = v0.2 + E5 自评回灌禁区 + runway 记账成本 + 纵向 Eval 运行拓扑 + 裁决角色分权。上位宪法:
docs/architecture/eval-philosophy.md(七公理 E0–E6;v2.2 已 ratified)。 机制选择与落地边界以 ADR-031 v3.4 为准:按问题选机制, 不按层补齐。
Why This Is a Skill(价值门禁)
模型的训练先验里有"怎么算 accuracy",没有猫咖的 eval 宪法。本 skill 是 2026-07 Eval 宪法(七公理 E0–E6)+ Alden 三日思辨的操作化——未来任何猫给任何 harness 域 配 eval 时,用同一套流程,防止家里的 eval 资产继续长成划水/污染/归因停滞/ 干预失证/摸鱼五病混合物。 来源:docs/architecture/eval-philosophy.md(宪法 v2.2,七公理)+ feature-discussions/2026-07-17-eval-charter-draft.md(v1 历史草案,已演进归档)+ 2026-07-16-alden-dialogue-distillation.md(思辨蒸馏)。
核心定义:一个 eval 指标到底是什么
一个 eval 指标 = 一个赌注:"这个数字的变动方向,与我们真实在乎的东西的变动 方向一致。" 它不是测量值本身,是"测量值 → 真实效用"的映射假设。假设会失效 (分布漂移)、会被破坏(优化压力)、需要验证(校准)——所以指标有生命周期。
〇、E0 资格门(建前先过——三问任一答不出即不发牌照,宪法 E0)
- claim 落在哪个 GT 域——同一产物的不同 claim 落在不同域(测试判得了 "符合规约",判不了"是人要的产品"),禁止拿产物整体贴单一域标签;
- 判定所需的新鲜 bit 在哪——已在系统可观察边界内(verifier 射程:边际 判分成本≈0、可自动重复)/ 在冻结先验里(judge 射程:便宜但有额度)/ 只在 价值主人脑中与真实后果里(必须外采)。裁判类型由 claim 的域决定,不由 预算决定;
- 裁判的工资谁在付——verifier=测试维护投入;judge=冻结先验额度 + calibration runway 的记账/人锚抽样成本(见出生证);价值主人=真实关系与打扰 预算。答不出付薪方=环转不久。若连 runway 都付不起,结论是缩小或不建 judge 环, 不是省掉测量后继续宣称它可持续。
E0 划的是自治上限:外部新 bit 缺位,不得宣称开放价值 claim 已全自动闭环 改善——但不否决候选生成、分诊、shadow、灰度等局部自动化。建前定资格,运行中 随 runway 持续重验。judge 联网算不算破圈 → 宪法 E0 联网三分判:查事实=真升级 / 查人群资料=换仓库的老本 / 接该用户本人历史=真破圈(破圈靠接上真值,不靠搜索动作)。
一、指标出生证契约(适用字段缺一不发牌照)
任何新指标上线前填齐基础字段;使用 judge 或其他会折旧、存在额度的裁判时, 再填额度字段;只有累计证据、持续运行的纵向 Eval 才填运行拓扑字段:
metric_birth_certificate:
utility_claim: # 这个数字上升,代表什么真实的东西变好?(答不出 = 拒发)
estimator: # 分子/分母/排除项/采样方式/judge 及其版本
validity_bounds: # 预注册失效条件:什么分布漂移/优化压力/judge 变化会让它失真
roles: # 四角色分别写明;允许兼任,但须在 role_overlap_justification 解释
observer: # 谁读 canonical evidence、生成测量
domain_owner: # 谁拥有该域规约与真相
consumer: # 谁据此做 keep/tune/sunset(无 consumer = 摸鱼指标,拒发)
calibrator: # 谁独立检查量尺/观察面可靠性
role_overlap_justification: # 同一主体兼任时,凭什么仍有独立性;开放价值 claim 从严
calibration_plan: # 多久和人工裁决/外生 ground truth 对一次表;相关性掉线阈值
repeatability_contract: # (v0.1 增,Sol 刀③)本指标属发现/归因/验收哪一环节;
# episode/环境/judge/版本如何冻结;跑几次;均值与 CI 波动
# 容差;哪些随机源允许变化。校准管"测得准",本件管"重测稳"。
# 以下两项仅对使用 judge / 其他有额度裁判的环适用;纯 verifier 环不填。
calibration_runway: # (v0.2 增,宪法 E0)适用时必填:额度不可读成单一电量,
# 按向量三账估——校准账(人-judge 决策级分歧率,晋升/回滚
# 被翻转才算)/ 暴露账(judge/题库被用于自适应选型的轮数)
# / 覆盖账(分布外流量占比)。抽样两腿:随机盲抽 + 风险
# 定向抽;复用用户自然行为,不逼打分。仪表 consumer=守门猫。
exhaustion_action: # (v0.2 增,宪法 E0)与 runway 同条件必填:额度告急的
# 预注册动作:自动晋升降
# shadow / 转人工复核 / 回退简单基线。分级停止判据:决策
# 分歧率越过预注册阈值=单独硬停(锚级);输出多样性坍缩=
# gaming 调查(统计级);打不赢最低充分替代物=经济退场。
# 以下块只对累计证据、持续运行的纵向 Eval 适用;单次交付检查不填。
longitudinal_trigger_contract:
trigger_policy: # event_plus_time | time_only;不把长期 Eval 配成无兜底 event_only
evidence_ingestion: # canonical episode 如何进入;入库成功不等于已经运行 Eval
early_trigger: # 可靠事件源下,什么阈值跨越/关键事件会提前唤醒
time_fallback: # 最长沉默多久必须复评;time_only 必须声明最大检测延迟
dedupe_key: # event 与 time 同时命中时如何归并成一次窗口
overlap_policy: # 上一轮仍运行时,第二次触发如何 queue/coalesce/reject
maturity_predicate: # 什么条件说明证据窗口已成熟,可计算有效测量
actionability_gate: # validity、权限、校准满足什么条件,verdict 才能驱动动作
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 211 lines · 90 tokens per session scan A 00a8dbaf3441
eval-design is a skill published in the GitHub repository zts212653/clowder-ai (2,894 stars, last pushed today), licensed MIT. It adds 90 tokens to every session and 4,715 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
local-ai-agents
Build local-first AI agents that run entirely on a developer workstation with Microsoft Foundry Local and Qwen function-calling models. Covers Small Language Models (SLMs), the OpenAI-compatible local endpoint, sandboxed local tools, local RAG with Chroma, local MCP servers, hybrid cloud/local routing, and the…
next-cache-components-adoption
Turn on Cache Components in a Next.js app and resolve the blocking routes it surfaces. Use when the user wants to enable, adopt, or migrate to Cache Components, flip the cacheComponents flag, work through a flood of blocking-prerender / instant validation errors, run the cache-components-instant-false codemod, or…
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…
insight-error-page
Write or audit an insight-kind error page for the Next.js dev overlay. Use when creating a new errors/ .mdx page, auditing an existing one, or checking that a page matches the framework fix cards. Covers page structure, title alignment, FixCard cards with Copy prompt button, code snippets, terminology verification…