eval-driven-development

eval-driven-development is a skill for Claude Code, Codex from ContextJet-ai/awesome-llm-observability. It costs 88 tokens per session (612 once invoked), scanned A, original, no licence file.

A development method for AI features that writes evaluations before making changes. An evaluation, or eval, is a repeatable test that checks how well a model or agent handles example tasks.

In plain words
What is it for?
Building and improving LLM features, prompts, models, and RAG systems through testable iterations.
Why use it?
It helps you tell whether a prompt, model, or retrieval change actually improved the feature and whether it broke existing behavior.

Skill for Claude CodeCodex

Part of the llm-observability plugin — 26 skills shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/contextjet-ai/awesome-llm-observability/eval-driven-development
Any agent
npx skills add ContextJet-ai/awesome-llm-observability --skill eval-driven-development
Clone the repo
git clone --depth 1 https://github.com/ContextJet-ai/awesome-llm-observability

Made for: Claude Code, Codex.

Or install llm-observability, the plugin that ships this one along with the rest of its 26 skills.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for eval-driven-development

README.md
[![agentmods](https://agentmods.dev/badge/skills/contextjet-ai/awesome-llm-observability/eval-driven-development.svg)](https://agentmods.dev/skills/contextjet-ai/awesome-llm-observability/eval-driven-development)
Your own site
<a href="https://agentmods.dev/skills/contextjet-ai/awesome-llm-observability/eval-driven-development"><img src="https://agentmods.dev/badge/skills/contextjet-ai/awesome-llm-observability/eval-driven-development.svg" alt="Measured on agentmods" height="20"></a>
Per session 88 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 612 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin unknown No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00088 $0.00612
Opus 5 $0.00044 $0.00306
Sonnet 5 $0.00018 $0.00122
Haiku 4.5 $0.00009 $0.00061

Measured 5d ago against content hash 245a0f873f0d, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

eval-driven-development scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/eval-driven-development/SKILL.md · 37 lines

The source is not reproduced here

A licence we could not identify

The repository carries a LICENSE file, but it is custom or dual enough that GitHub cannot name it and neither can this catalogue. Unknown terms are not permission, so the body is not copied here. Read the licence at the source and decide for yourself.

Read it on GitHub

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 5d ago First seen · 37 lines · 88 tokens per session scan A 245a0f873f0d

Subscribe to this mod's changes

eval-driven-development is a skill published in the GitHub repository ContextJet-ai/awesome-llm-observability (33 stars, last pushed 4d ago), with no licence file. It adds 88 tokens to every session and 612 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

kling-prompter

生成符合 Kling 3.0(可灵 3.0,快手)规则的视频提示词。三种写法自适应:4 部分基础公式(短视频)/ 5 层进阶公式(剧情+音频)/ 图生视频专用(只描述运动)。Kling 是 2026 年中文理解最强、原生音画同步、最长 2 分钟、支持角色定向发声、Motion Brush 的电影级模型。用于"用 Kling 生成视频"、"可灵 AI 提示词"、"中文视频生成"、"带原生音频的剧情视频"、"图生视频"、"角色对话视频"等触发场景。.

cclank/lanshu-awesome-ai-video-kit · 162 tokens

happyhorse-prompter

生成符合 HappyHorse 1.0 严格规则的紧凑提示词(30-55 词),主体先行 + 明确镜头技术 + 音频激活路径(with X audible / speaking English at natural pace)。可选加入"8s 时序节拍"结构。用于"用 HappyHorse 生成视频"、"做个 3-15 秒短片"、"要原生带音频的视频"、"ASMR 视频提示词"等触发场景。.

cclank/lanshu-awesome-ai-video-kit · 116 tokens

seedance-prompter

把用户的自然语言视频需求转换为符合 Doubao Seedance 2.0 进阶公式的提示词(8 要素:精准主体+动作细节+场景环境+光影色调+镜头运镜+视觉风格+画质+约束条件)。用于"帮我写一个 Seedance 提示词"、"生成视频提示词"、"做个产品广告视频"、"用 Seedance 生成 XX"等触发场景。如果用户没明确说 Seedance,但描述的是单一镜头的复杂叙事/多主体/电影感视频,也优先用此 skill。.

cclank/lanshu-awesome-ai-video-kit · 142 tokens

claude-api

Anthropic Claude API patterns for Python and TypeScript. Covers Messages API, streaming, tool use, vision, extended thinking, batches, prompt caching, and Claude Agent SDK. Use when building applications with the Claude API or Anthropic SDKs.

hashgraph-online/awesome-codex-plugins · 53 tokens

iterative-retrieval

Pattern for progressively refining context retrieval to solve the subagent context problem.

hashgraph-online/awesome-codex-plugins · 19 tokens

cost-aware-llm-pipeline

成本感知 LLM 管道,根据任务复杂度选择合适模型, 管理上下文预算,避免在长会话末尾做大型重构。.

hashgraph-online/awesome-codex-plugins · 42 tokens