evaluator-development

A guide for adding evaluators to a software project, including rule-based checks, language-model checks, and agent-based checks. An evaluator is a component that examines input and returns an assessment.

In plain words
What is it for?
Use it when developing evaluators for default checks, training data, benchmarks, supervised fine-tuning, retrieval-augmented generation, or hallucination detection.
Why use it?
It explains the project structure, base classes, registration methods, and run groups needed to add evaluators consistently.

Cursor rule for Cursor

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add rules/migoxlab/dingo/evaluator-development
Clone the repo
git clone --depth 1 https://github.com/MigoXLab/dingo

Made for: Cursor.

Per session 0 Nothing until a file matches its globs; then the whole rule loads.
When invoked 791 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00000 $0.00791
Opus 5 $0.00000 $0.00396
Sonnet 5 $0.00000 $0.00158
Haiku 4.5 $0.00000 $0.00079

Measured yesterday against content hash e9a39b0f5e73, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

evaluator-development scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.cursor/rules/evaluator-development.mdc · 99 lines

How it starts

The opening of the file, as written. The whole thing — 99 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Evaluator Development Guide

Architecture

Model (registry)
├── Rule evaluators  → @Model.rule_register(metric_type, groups)
│   └── BaseRule.eval(input_data) → EvalDetail
├── LLM evaluators   → @Model.llm_register(name)
│   └── BaseOpenAI.eval(input_data) → EvalDetail
└── Agent evaluators → @Model.llm_register(name)
    └── BaseAgent.eval(input_data) → EvalDetail

Key Files

File Purpose
dingo/model/model.py Model class with rule_register, llm_register, load_model
dingo/model/rule/base.py BaseRule base class
dingo/model/llm/base_openai.py BaseOpenAI base class for LLM evaluators
dingo/model/llm/agent/base_agent.py BaseAgent base class for agent evaluators
dingo/io/input/data.py Data model (input to evaluators)
dingo/io/output/eval_detail.py EvalDetail model (output from evaluators)

Registration Groups

Rules belong to groups that determine when they run:

  • default — runs in default evaluation
  • pretrain — pre-training data quality
  • benchmark — benchmark evaluation
  • sft — supervised fine-tuning data
  • rag — RAG system evaluation
  • hallucination — hallucination detection

EvalDetail Contract

Every evaluator must return EvalDetail with:

Field Type Description
metric str Evaluator class name (cls.__name__)
status bool True = issue found, False = no issue
label List[str] Quality labels (e.g., ['QUALITY_GOOD'] or ['QUALITY_BAD_COMPLETENESS.RuleName'])
reason List[str] Human-readable explanation of the finding
score float Optional numeric score (0.0–1.0)
extra Dict Optional extra metadata

Data Field Access

Since Data uses extra = "allow", always access non-standard fields safely:

# Safe access patterns
raw_data = getattr(input_data, 'raw_data', {})
context = getattr(input_data, 'context', None)
reference = getattr(input_data, 'reference', '')

# For RAG evaluators, common field access pattern
question = input_data.prompt or raw_data.get("question", "")
answer = input_data.content or raw_data.get("answer", "")

Read the full file on GitHub · 99 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 99 lines · 0 tokens per session scan A e9a39b0f5e73

Subscribe to this mod's changes

evaluator-development is a cursor rule published in the GitHub repository MigoXLab/dingo (751 stars, last pushed 4d ago), licensed Apache-2.0. It costs nothing until one of its globs matches a file; then it loads 791 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.