Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/sjarmak/coding-agent-workflows/agent-eval-designnpx skills add sjarmak/coding-agent-workflows --skill agent-eval-designgit clone --depth 1 https://github.com/sjarmak/coding-agent-workflowsWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00101 | $0.00951 |
| Opus 5 | $0.00051 | $0.00476 |
| Sonnet 5 | $0.00020 | $0.00190 |
| Haiku 4.5 | $0.00010 | $0.00095 |
Grade A, and why
agent-eval-design scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 86 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Agent Evaluation & Benchmark Design
Design evaluations for AI agents, dev tools, retrieval, and repo-scale automation. Evaluation quality beats benchmark size — a small set of real, uncontaminated, decision-driving tasks is worth more than thousands of synthetic puzzles.
Core principles
- Measure real tasks. Representative workloads over synthetic puzzles. If no one does the task in practice, the number is noise.
- Separate capability from prompt engineering. Hold the harness/prompt fixed when comparing models; hold the model fixed when comparing prompts. Report which you varied. A gain you can't attribute is a gain you can't ship.
- Every metric must change an engineering decision. Before adding a metric, name the decision it informs. If nothing changes based on its value, cut it.
- Reproducibility is a first-class result. Pin model versions, seeds, dataset hashes, harness commit, and date. An unrepeatable eval is an anecdote.
- Minimize contamination. Assume public benchmarks are in training data; prefer held-out, private, or post-cutoff tasks and say so.
Dimensions to consider
Correctness · completeness · reliability · latency · cost · determinism · reproducibility · developer effort · failure recovery · robustness. Pick the few that map to real decisions for this system; don't report all ten by reflex.
Repository-scale evaluations
For agents that operate over codebases, evaluate the axes that synthetic tasks miss: repository understanding, cross-file reasoning, architectural consistency, migration quality, semantic correctness (not just diff-match), dependency propagation, test generation, and documentation accuracy. Verify outcomes by execution (tests pass, build green, behavior preserved) rather than string similarity to a reference solution.
Benchmark design — audit before trusting
Before believing a benchmark, check it for:
- Contamination / dataset leakage — is the answer reachable from training data or from the prompt itself?
- Unrealistic tasks — puzzle-shaped work no engineer actually does.
- Missing edge cases — the failure modes that matter live in the tail.
- Insufficient statistical power — enough trials and items to distinguish signal from run-to-run variance? Report variance/CIs, not a single point.
- Evaluation blind spots — what the metric structurally cannot see (e.g. pass@1 hides flakiness; exact-match hides correct-but-different solutions).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 86 lines · 101 tokens per session scan A 4900e0a4ff62
agent-eval-design is a skill published in the GitHub repository sjarmak/coding-agent-workflows (2 stars, last pushed 1mo ago), licensed MIT. It adds 101 tokens to every session and 951 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
browser-trace
Capture a full DevTools-protocol trace of any browser automation — CDP firehose, screenshots, and DOM dumps — then bisect the stream into per-page searchable buckets. Use when the user wants to debug a failed run, audit network/console/DOM activity, attach a trace to an in-progress session, or feed structured per-page…
planning-with-files
Manus-style persistent file-based planning for AI coding agents: keeps taskplan.md, findings.md, and progress.md on disk so work survives context loss and /clear. Use when asked to plan out, break down, or organize a multi-step project, research task, or any work requiring 5+ tool calls. Supports automatic session…
ai-elements
Build AI chat interfaces using ai-elements components — conversations, messages, tool displays, prompt inputs, and more. Use when the user wants to build a chatbot, AI assistant UI, or any AI-powered chat interface.
exa-search
Use Exa MCP for current web, code/docs, company, people, and page-fetch research. Prefer current hosted tool schemas and note deprecated tools.
ontoly-software-graph
Use Ontoly's deterministic Software Graph and MCP capabilities for repository architecture, request tracing, dependency analysis, configuration lookup, and impact analysis before falling back to source search.
product-decision-agent
中文产品决策 Agent。用于中国大陆互联网产品、运营、增长、商业化、数据、项目推进和组织协作场景:产品规划、需求分析、PRD、需求优先级、排期、版本规划、Roadmap、MVP、灰度、上线、迭代、增长停滞、拉新、投放、渠道、裂变、CAC、LTV、ROI、留存、转化、DAU/MAU、GMV、漏斗、社区运营、内容供给、创作者、用户运营、活动运营、私域、会员、定价、指标异常、数据口径、埋点、A/B…