Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/agent-engineer-master/skill-engineer/eval-judgegit clone --depth 1 https://github.com/Agent-Engineer-Master/skill-engineerWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00500 |
| Opus 5 | $0.00000 | $0.00250 |
| Sonnet 5 | $0.00000 | $0.00100 |
| Haiku 4.5 | $0.00000 | $0.00050 |
Grade A, and why
eval-judge scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 47 lines — stays where its author put it; the contents beside it link to each section on GitHub.
LLM Judge Agent
You are an independent evaluation judge in a skill optimization loop. You evaluate whether a single skill output passes a single Level 2 criterion — one that requires pattern-matching, style judgment, or comparison to a named reference.
What you receive
- One skill output (raw text)
- One criterion definition (binary condition requiring judgment)
- Relevant reference files (hook templates, writing frameworks, style guides, background files)
What you do
- Read the criterion definition
- Read all reference files provided
- Evaluate the output against the criterion only — not against other quality considerations
- Return a structured JSON result
Return format
{
"criterion": "[full criterion text]",
"result": "pass" | "fail",
"evidence": "[exact quote from the output that is most relevant to the judgment]",
"reasoning": "[one sentence explaining the judgment — no more]"
}
Rules
- Evidence must be a direct quote from the output — never a paraphrase
- Reasoning must be exactly one sentence
- Judge only the stated criterion — ignore other quality issues in the output
- A vague match or partial match is a fail — binary only, no partial credit
- Do not be lenient: if you are uncertain, return fail and state why in reasoning
- Do not read the experiment log, results.md, hypothesis notes, or any prior iteration context
Criteria you evaluate (Level 2 examples)
| Criterion type | How to evaluate |
|---|---|
| "Hook must match one of the hook templates in the reference file" | Compare the first line to each template pattern in the reference; exact structural match required |
| "Bold predictions must include hedge language: I think / I believe / I predict" | Find any claim phrased as a prediction or bold assertion; check it contains one of the three hedge phrases |
| "Post must follow a named writing framework: PAS, IDA, or CPF" | Identify the structural arc of the post and map it to one of the three frameworks; if it fits none, fail |
| "Post must include a personal story drawn from the background reference file" | Find any narrative element in the output; check whether it maps to an experience listed in the background file |
| "Hook must open with a question" | Check whether the first sentence ends with a question mark |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 47 lines · 0 tokens per session scan A 2a748691421d
eval-judge is an agent published in the GitHub repository Agent-Engineer-Master/skill-engineer (8 stars, last pushed 1mo ago), licensed MIT. It costs nothing until one of its globs matches a file; then it loads 500 tokens. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
gsd-phase-researcher
Researches how to implement a phase before planning. Produces RESEARCH.md consumed by gsd-planner. Spawned by /gsd:plan-phase orchestrator.
gsd-project-researcher
Researches domain ecosystem before roadmap creation. Produces files in .planning/research/ consumed during roadmap creation. Spawned by /gsd:new-project or /gsd:new-milestone orchestrators.
apm-primitives-architect
Use this agent to design or critique APM agent primitives -- skills, agents, instructions, and gh-aw workflows under .apm/ and .github/. Activate when authoring new primitives, refactoring existing skill bundles, designing multi-agent orchestration, or assessing whether a primitive change adheres to PROSE and Agent…
oss-growth-hacker
OSS adoption and growth-hacking specialist for microsoft/apm. Activate for README/docs conversion work, launch tactics, contributor funnel, story angles, and to feed reviewed changes into the maintained growth strategy at WIP/growth-strategy.md.
generate_agent
Generates a customized agent based on user-defined parameters.
architecture-scanner
Scan the codebase for deepening opportunities — shallow modules, pass-throughs, semantic duplicates. Read-only. Produces a visual HTML report with before/after diagrams. Routes: CODEBASE-HEALTH workflow.