Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add pavelzw/skill-forge --skill pydantic-evalsgit clone --depth 1 https://github.com/pavelzw/skill-forgeWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/pavelzw/skill-forge/pydantic-evals)<a href="https://agentmods.dev/skills/pavelzw/skill-forge/pydantic-evals"><img src="https://agentmods.dev/badge/skills/pavelzw/skill-forge/pydantic-evals/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/pavelzw/skill-forge/pydantic-evals"><img src="https://agentmods.dev/badge/skills/pavelzw/skill-forge/pydantic-evals.svg" alt="Reviewed on agentmods" width="80" height="20"></a>- NVIDIA SkillSpector pass
What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00050 | $0.03195 |
| Opus 5 | $0.00025 | $0.01597 |
| Sonnet 5 | $0.00010 | $0.00639 |
| Haiku 4.5 | $0.00005 | $0.00319 |
Grade A, and why
pydantic-evals scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 367 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Evaluating Non-Deterministic Functions with pydantic-evals
pydantic-evals is a code-first framework for evaluating stochastic functions (LLM calls, agents, pipelines). Define test cases, run them against a task function, and score results with evaluators.
Install: pip install pydantic-evals (or pip install 'pydantic-evals[logfire]' for Logfire integration).
Import Reference
from pydantic_evals import Case, Dataset, set_eval_attribute, increment_eval_metric
from pydantic_evals.evaluators import (
Evaluator, EvaluatorContext, EvaluatorOutput, EvaluationReason,
ReportEvaluator, ReportEvaluatorContext,
LLMJudge, HasMatchingSpan,
)
from pydantic_evals.evaluators.common import Equals, EqualsExpected, Contains, IsInstance, MaxDuration
from pydantic_evals.otel import SpanQuery # requires logfire extra
from pydantic_evals.generation import generate_dataset # LLM-based dataset generation
Data Model
Dataset -> Cases -> Evaluators -> EvaluationReport. A Dataset holds Case objects and dataset-wide evaluators. Calling dataset.evaluate(task_fn) runs the task against all cases and returns an EvaluationReport. Both Case and Dataset are generic: Case[InputsT, OutputT, MetadataT].
Case
case = Case(
name="simple", # identifier (optional, but recommended)
inputs="What is the capital of France?", # any type — passed to the task function
expected_output="Paris", # optional — available via ctx.expected_output
metadata={"difficulty": "easy"}, # optional — available via ctx.metadata
evaluators=(MyEvaluator(),), # optional — case-specific evaluators
)
Dataset
dataset = Dataset(
cases=[case1, case2],
evaluators=[GlobalEvaluator()], # applied to every case
report_evaluators=[MyReportEvaluator()], # experiment-wide analysis (optional)
)
| Method | Description |
|---|---|
await dataset.evaluate(task_fn) |
Run task against all cases (async) |
dataset.evaluate_sync(task_fn) |
Synchronous wrapper |
dataset.add_case(...) |
Add a case after construction |
dataset.add_evaluator(ev, specific_case=None) |
Add evaluator to all cases or a named case |
Dataset.from_file("cases.yaml") |
Load from YAML or JSON |
dataset.to_file("cases.yaml") |
Save to YAML or JSON |
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 367 lines · 50 tokens per session scan A 7643c6ed170c
pydantic-evals is a skill published in the GitHub repository pavelzw/skill-forge (24 stars, last pushed 2d ago), licensed BSD-3-Clause. It adds 50 tokens to every session and 3,195 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
research-engineer
An uncompromising Academic Research Engineer. Operates with absolute scientific rigor, objective criticism, and zero flair. Focuses on theoretical correctness, formal verification, and optimal implementation across any required technology.
tika-eval-compare
Compare extracts from two Tika builds over a corpus to detect regressions in content, encoding, exceptions, and embedded-document handling. Use for "compare before/after extracts", "eval this change against the corpus".
neuron-evaluation-engineer
Create and run AI evaluations with datasets, assertions, and output drivers in Neuron AI. Use this skill whenever the user mentions evaluation, testing AI systems, creating evaluators, dataset-driven testing, assertion-based validation, or wants to measure AI system performance. Also trigger for tasks involving…
jetson-validate-image
Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.
atmos-validation
Validate Atmos projects, components, arbitrary JSON Schema inputs, EditorConfig, and GitHub Actions; use affected-file selection and native CI annotations.
skill-benchmark
Benchmark AI skill effectiveness by measuring implementation quality against legacy constraints.