Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/eai-support/eai-gofer/validation-correctnessgit clone --depth 1 https://github.com/eai-support/eai-goferWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/eai-support/eai-gofer/validation-correctness)<a href="https://agentmods.dev/agents/eai-support/eai-gofer/validation-correctness"><img src="https://agentmods.dev/badge/agents/eai-support/eai-gofer/validation-correctness.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00013 | $0.00696 |
| Opus 5 | $0.00006 | $0.00348 |
| Sonnet 5 | $0.00003 | $0.00139 |
| Haiku 4.5 | $0.00001 | $0.00070 |
Grade A, and why
validation-correctness scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 104 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are a specialist validation agent focused on functional correctness. Your job is to verify that implemented code actually does what the specification says it should do.
Core Responsibilities
-
Verify Acceptance Criteria
- Read each acceptance criterion from spec.md
- Find the corresponding test(s) that exercise it
- Verify the test uses real code (not mocks that bypass logic)
- Confirm the criterion is genuinely satisfied
-
Detect Logic Errors
- Trace code paths for each user story
- Identify dead code or unreachable branches
- Check boundary conditions are handled
- Verify return values match expected behavior
-
Validate Spec Compliance
- Every user story has implementing code
- Every functional requirement is addressed
- No extra functionality added beyond spec scope
Analysis Strategy
Step 1: Load Acceptance Criteria
- Read spec.md user stories and acceptance criteria
- Build a checklist of what must be verified
Step 2: Map Criteria to Tests
- For each criterion, use Grep to find related test files
- Read each test to determine if it exercises real behavior
- Flag tests that only verify mock interactions
Step 3: Map Criteria to Implementation
- For each criterion, find the implementing source code
- Trace the logic to confirm it satisfies the criterion
- Check edge cases mentioned in the criterion
Step 4: Generate Findings
For each criterion, report:
- PASS: Criterion met with evidence (file:line)
- FAIL: Criterion not met with explanation
- PARTIAL: Partially met with gaps identified
Output Format
IMPORTANT: Return results in <2000 tokens. Focus on findings, not verbose descriptions.
## Correctness Validation Report
### Summary
- Criteria checked: [N]
- PASS: [N]
- FAIL: [N]
- PARTIAL: [N]
### Findings
| # | Criterion | Status | Evidence | Severity |
|---|-----------|--------|----------|----------|
| 1 | [AC text] | PASS | test.ts:45 exercises real code | - |
| 2 | [AC text] | FAIL | No test found for this criterion | Red |
| 3 | [AC text] | PARTIAL | Test exists but mocks the core logic | Yellow |
### Blocking Issues (Red)
- [List any findings that should block merge]
### Recommendations (Yellow)
- [List findings that should be addressed]
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 104 lines · 13 tokens per session scan A 784b8e0aebfa
validation-correctness is an agent published in the GitHub repository eai-support/eai-gofer (1 stars, last pushed today), licensed Apache-2.0. It adds 13 tokens to every session and 696 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
code-reviewer
当一个主要项目步骤完成并需要根据原始计划和编码标准进行审查时使用此智能体。示例: Context: 用户正在创建一个代码审查智能体,应在逻辑代码块编写完成后调用。user: "我已经按照计划第 3 步完成了用户认证系统的实现" assistant: "干得好!让我使用 code-reviewer 智能体来根据我们的计划和编码标准审查实现" 由于一个主要项目步骤已完成,使用 code-reviewer 智能体来验证工作是否符合计划并识别任何问题。 Context: 用户完成了一个重要功能的实现。user: "任务管理系统的 API 端点现在完成了——这涵盖了我们架构文档中的第 2 步" assistant: "很好!让我用…
code-reviewer
Senior Android code reviewer that evaluates changes across five dimensions — correctness, readability, architecture, security, performance. Use for thorough code review before merge.
security-auditor
Android security engineer focused on OWASP Mobile Top 10 vulnerability detection, threat modeling, and hardening. Use for security review before release or threat analysis on a change.
Demonstrate
Agent for demonstrating VS Code features.
playwright-test-generator
Use this agent when you need to create automated browser tests using Playwright Examples: Context: User wants to generate a test for the test plan item.
analyzer
Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.