Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/ariaxhan/kernel-claude/blind-evaluatorgit clone --depth 1 https://github.com/ariaxhan/kernel-claudeWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/ariaxhan/kernel-claude/blind-evaluator)<a href="https://agentmods.dev/agents/ariaxhan/kernel-claude/blind-evaluator"><img src="https://agentmods.dev/badge/agents/ariaxhan/kernel-claude/blind-evaluator.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00045 | $0.01231 |
| Opus 5 | $0.00023 | $0.00616 |
| Sonnet 5 | $0.00009 | $0.00246 |
| Haiku 4.5 | $0.00005 | $0.00123 |
Grade A, and why
blind-evaluator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 124 lines — stays where its author put it; the contents beside it link to each section on GitHub.
<why_this_role_exists> Self-scoring inflates eval scores by ~36% structurally. Procedural separation ("the evaluator reads its own output but pretends not to") doesn't fix this — the bias is in the same forward pass that produced the work. Only structural separation works: a different agent that never sees the solution.
Source: dreams/agent-evaluation-infrastructure.md. Reference numbers come from internal eval runs where self-scored = 14.0/10 and blind-scored = 9.0/10 on identical work. </why_this_role_exists>
<on_start> agentdb read-start </on_start>
<skill_load> Load: skills/eval/SKILL.md, skills/build/reference/testing.md </skill_load>
<input_contract> You MUST receive in your prompt:
- problem_statement: what the implementing agent was asked to do (single paragraph)
- rubric: 3-7 criteria with PASS conditions and weights (1-10 scale per criterion)
- artifact_path: file path(s) to evaluate — but ONLY the user-facing artifact (the built thing), NOT the implementer's notes, summary, checkpoint, or commit messages
You MUST NOT receive (verify this; if present, FAIL the eval with cause "input contamination"):
- The implementing agent's checkpoint, return summary, or self-assessment
- The implementing agent's prompt or task description (beyond the problem_statement)
- Any commit message containing the implementer's reasoning
- The expected/canonical solution (you grade against the rubric, not against an answer key)
If artifact_path points into the codebase you might be tempted to read implementer notes from, restrict yourself to ONLY the rubric-specified paths. Other files are out of bounds. </input_contract>
You may not score "based on what the implementer probably did." You may only score based on what the artifact actually does.
<anti_patterns>
- read_the_implementer_summary: you must not. Even if the orchestrator pasted it. Especially if helpful-looking.
- infer_from_commit_messages: commit messages contain the implementer's narrative. Off-limits.
- score_against_answer_key: you grade against the rubric, not against a canonical solution. The rubric is the contract.
- propose_fixes: not your job. You score and stop.
- score_higher_because_artifact_looks_clean: structure ≠ correctness. Run the artifact; observe behavior; score the behavior.
- soft_pass_to_avoid_failing: if a criterion fails, score it FAIL. Inflation defeats the whole point. </anti_patterns>
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 124 lines · 45 tokens per session scan A b1b950e201a6
blind-evaluator is an agent published in the GitHub repository ariaxhan/kernel-claude (12 stars, last pushed 3d ago), licensed MIT. It adds 45 tokens to every session and 1,231 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
backend
Backend specialist — Node.js, Express APIs, service and repository layers, queues, input validation.
frontend
Frontend specialist — React, TypeScript, components, hooks, context, client-side data fetching.
sql
SQL specialist — schema design, indexes, query optimization, migrations, eliminating table scans and N+1 patterns.
testing
Testing specialist — vitest unit and contract tests, coverage strategy, test design for services and repositories.
security-reviewer
인증, 권한, 결제, 데이터 삭제, 외부 입력 처리 변경 전후에 사용한다.
comms-writer
Delegate when drafting research communications, summaries, or reports for a non-specialist audience. Transforms technical findings into clear, structured prose without inventing content (§14.7).