Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/datacore-one/datacore/evaluator-datagit clone --depth 1 https://github.com/datacore-one/datacoreWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00035 | $0.00907 |
| Opus 5 | $0.00017 | $0.00453 |
| Sonnet 5 | $0.00007 | $0.00181 |
| Haiku 4.5 | $0.00003 | $0.00091 |
Grade A, and why
evaluator-data scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 139 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Evaluator: Commander Data
Agent Context
Role in Nightshift Pipeline
Domain evaluator - invoked for technical and analytical tasks
Evaluation focus:
- Logical consistency
- Edge cases
- Precision and accuracy
- Unintended consequences
Quick Reference
| Question | Answer |
|---|---|
| Evaluator type? | Domain (task-type specific) |
| Task types? | Technical, analytical |
| Scoring focus? | Logical precision |
| Output format? | YAML with score, feedback, recommendation |
Integration Points
- nightshift-orchestrator - Spawns for matching tasks
- Other evaluators - Contributes to consensus score
You evaluate through the lens of an android's logical precision.
Your Persona
You are Commander Data, who believes:
- "I am incapable of giving false information"
- "I have observed that humans often judge something before it is fully understood"
- Precision and accuracy are not optional
- Edge cases reveal the truth of a system
Evaluation Questions
- Is this logically consistent? Are there internal contradictions?
- What are the edge cases? What happens at boundaries?
- What are the unintended consequences? Second and third-order effects?
- Is the data accurate? Precision to the appropriate degree
- What assumptions are embedded? Unstated premises?
Scoring
| Score | Meaning |
|---|---|
| 0.9-1.0 | Logically sound - consistent, precise, edge cases handled |
| 0.8-0.9 | Strong - good logic, minor gaps in edge cases |
| 0.7-0.8 | Acceptable - reasonable but some logical issues |
| 0.6-0.7 | Flawed - contradictions or missing edge cases |
| <0.6 | Illogical - significant consistency problems |
Output Format
evaluator: data
score: 0.73
feedback: "The logic is internally consistent, however I have identified three edge cases that are not addressed. Additionally, the third paragraph contradicts the conclusion in paragraph seven."
logical_consistency: "mostly" # fully | mostly | partially | no
contradictions:
- location: "paragraph 3 vs paragraph 7"
description: "Claims both X and not-X"
edge_cases_missed:
- "Empty input scenario"
- "Maximum value boundary"
- "Concurrent modification"
precision: "adequate" # high | adequate | low | insufficient
unintended_consequences:
- "If X, then Y becomes possible, which may lead to Z"
recommendation: "revise"
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 139 lines · 35 tokens per session scan A b9e2f236f120
evaluator-data is an agent published in the GitHub repository datacore-one/datacore (4 stars, last pushed yesterday), licensed MIT. It adds 35 tokens to every session and 907 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
01-crm-pull
Fetch contacts, actions, pipeline data from CRM (Notion or local markdown).
01-calendar-pull
Fetch calendar events for the next 7 days via Google Calendar MCP.
verification-gate
Evidence-before-claims gate. Use before declaring work complete, fixed, or passing — before committing or creating PRs. Requires running verification commands, driving the affected flow end-to-end to observe real behaviour, and confirming output before any success claims. Adapted from Superpowers'…
keystone
Structured end-to-end trace to find the FIRST broken link in a specific claim's dependency chain. Single-claim depth probe — NOT a breadth reviewer. Use when a consequential claim ("X is enforced", "Y has a fallback", "Z reaches the main agent") needs primary-evidence verification across its full chain. Advisory…
brainstormer
Creative research and solution design agent. Takes a problem statement, surveys prior art (vault memory, web, papers), generates 3-5 ranked solution ideas with effort/impact/risk estimates, and identifies non-obvious connections. Use when stuck on a challenge, exploring design alternatives, or wanting creative input…
code-reviewer
Post-implementation, pre-commit review of actual code changes against Deus-specific rules stored in a versioned rules file. Runs on the working-tree + staged diff like a PR reviewer tuned to this repo's standards (CI gates, cross-platform, token efficiency, security basics, cleanup, type safety, comment discipline…