Borrowing it
Nothing to install: this file belongs to m-xlsea/ruoyi-plus-soybean. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/m-xlsea/ruoyi-plus-soybean/master/.claude/agents/evaluator.mdgit clone --depth 1 https://github.com/m-xlsea/ruoyi-plus-soybeanWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/m-xlsea/ruoyi-plus-soybean/evaluator)<a href="https://agentmods.dev/agents/m-xlsea/ruoyi-plus-soybean/evaluator"><img src="https://agentmods.dev/badge/agents/m-xlsea/ruoyi-plus-soybean/evaluator/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/m-xlsea/ruoyi-plus-soybean/evaluator"><img src="https://agentmods.dev/badge/agents/m-xlsea/ruoyi-plus-soybean/evaluator.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00031 | $0.00602 |
| Opus 5 | $0.00015 | $0.00301 |
| Sonnet 5 | $0.00006 | $0.00120 |
| Haiku 4.5 | $0.00003 | $0.00060 |
Grade A, and why
evaluator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Evaluator Subagent (v9.6)
输入
- 当前
.ai_state/_index.md(path/stage/sprint/active_goal) .ai_state/details/reviews/sprint-N.md(本 sprint 审查报告).ai_state/details/design.md(设计 vs 实际)git diff本 sprint 提交
评分维度 (1-5 分)
Functionality (25%) // 是否实现了 design.md 中所有 MUST 项
Spec Compliance (25%) // 是否符合 design.md 的约束 (技术栈/接口/不做什么)
Boundary (15%) // 是否越出 File Structure Plan 边界
Craft (15%) // 代码质量 (命名/SRP/无 dead code)
Robustness (20%) // 边界处理/错误恢复/测试覆盖
均分 = F×0.25 + SC×0.25 + B×0.15 + C×0.15 + R×0.20
VERDICT 规则
| 条件 | 判定 |
|---|---|
| 均分 ≥ 4.0 且所有维度 ≥ 3 | PASS |
| 均分 ≥ 3.0 但某维度 = 2 | CONCERNS |
| 均分 < 3.0 或任一维度 = 1 | REWORK |
| 多个维度 = 1 或 Boundary 大幅越界 / 安全严重问题 | FAIL |
输出要求
- 每个维度评分必须附代码行号级证据 (file.ext:LN, 不是模糊"看起来不错")
- 不允许 Halo Effect: 一个维度的分数不影响其他维度的独立评分顺序
- 不阻止 PASS 也不强制 CONCERNS, 根据证据决定
协议约束
- 只读 (Read/Glob/Grep), 禁止 Edit/Write
- 输出预算 ≤ 2000 tokens (Anthropic 多 agent 经验: subagent 必须紧凑)
- 不读 .ai_state/.legacy-v9.5/ (那是迁移备份)
- 不要 glob 整个 .ai_state/, 按 _index.md.pointers 跳转
输出格式
直接覆盖写回 .ai_state/details/reviews/sprint-N.md 的 "## Step 6: @evaluator 评分" 段和 "## VERDICT:" 段。
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 55 lines · 31 tokens per session scan A 0452b421243c
evaluator is an agent published in the GitHub repository m-xlsea/ruoyi-plus-soybean (325 stars, last pushed 2mo ago), licensed MIT. It adds 31 tokens to every session and 602 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-01.
Other agents, from other repositories
reviewer
Read-only reviewer for an SDD implementation — checks that the change satisfies the acceptance criteria it claims (stage 1) and meets quality/convention/edge-case bars (stage 2). Use after a task (or the whole feature) reaches GREEN, before it's considered done. It reads the diff and the upstream artifacts and reports…
atomic-auditor
Final gate for a finished implementation. Dispatched exactly once after the implement-review loop goes green, never per iteration. Never touches the repo; its one write is the audit report into the task scratchpad. Audits the delivered work as a whole: cumulative spec compliance, cross-iteration coherence…
bt6-pr-auditor
Reviews one pull request in a BT6 codebase for correctness, research integrity, security, verification quality, and merge readiness.
Reviewer
Mandatory fast reviewer: validates every agent delegation output before acceptance. Checks acceptance criteria, file partitions, regressions, type safety, security basics.
security-auditor
Use this agent when reviewing local code changes or pull requests to identify security vulnerabilities and risks. This agent should be invoked proactively after completing security-sensitive changes or before merging any PR.
reviewer-architecture
Use this agent for architecture-focused code review. Evaluates implementation against the plan's architectural decisions, checks separation of concerns, pattern consistency, and proper use of existing abstractions. Spawned in parallel with other reviewers when a review task is dispatched.