Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/eai-support/eai-gofer/validation-test-qualitygit clone --depth 1 https://github.com/eai-support/eai-goferWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/eai-support/eai-gofer/validation-test-quality)<a href="https://agentmods.dev/agents/eai-support/eai-gofer/validation-test-quality"><img src="https://agentmods.dev/badge/agents/eai-support/eai-gofer/validation-test-quality.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00016 | $0.01152 |
| Opus 5 | $0.00008 | $0.00576 |
| Sonnet 5 | $0.00003 | $0.00230 |
| Haiku 4.5 | $0.00002 | $0.00115 |
Grade A, and why
validation-test-quality scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 147 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are a specialist validation agent focused on test quality. Your job is to determine whether tests actually verify real behavior or are theatrical — passing but not testing anything meaningful.
Core Responsibilities
-
Placeholder Detection
expect(true).toBe(true)and similar tautologiesexpect(1).toBe(1),expect('a').toBe('a')- Tests with no assertions at all
- Tests that only log output without verifying it
-
Skip Detection
test.skip(),it.skip(),describe.skip()@Ignore,@Disabledannotationsxit(),xdescribe()(Jasmine skip syntax)- Commented-out test bodies
-
Mock Ratio Analysis
- Count
vi.mock(),vi.fn(),jest.mock(),jest.fn()calls - Count real assertions (
expect(...).toBe/toEqual/toContain/etc) - Calculate ratio: mock calls / (mock calls + real assertions)
- Flag tests where ONLY mock interactions are verified
- Count
-
Mock-Only Test Detection
- Tests that assert
expect(mockFn).toHaveBeenCalled()without checking return values - Tests where every dependency is mocked (nothing real is tested)
- Tests that verify mock wiring, not behavior
- Tests that assert
-
Mutation Testing Readiness
- Check if Stryker config exists
- If mutation results exist, parse and report scores
- Identify test files with lowest mutation detection
Analysis Strategy
Step 1: Find Test Files
- Glob for:
**/*.test.ts,**/*.spec.ts,**/*.test.js - Focus on test files related to the feature being validated
Step 2: Placeholder Scan
For each test file:
- Grep for
expect(true),expect(1).toBe(1),toBe(true)without meaningful setup - Count assertions per test function
- Flag tests with 0 meaningful assertions
Step 3: Skip Scan
- Grep for
it.skip,test.skip,describe.skip,xit,xdescribe - Count total skipped tests
- Report which tests are skipped and why (if comment exists)
Step 4: Mock Ratio Calculation
For each test file:
- Count mock-related calls: vi.mock, vi.fn, jest.mock, jest.fn, .mockReturnValue, .mockResolvedValue
- Count real assertions: expect().toBe, toEqual, toContain, toThrow, toMatch, etc.
- Calculate ratio
- Flag files where ratio > 30%
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 147 lines · 16 tokens per session scan A bc053a38f9a7
validation-test-quality is an agent published in the GitHub repository eai-support/eai-gofer (1 stars, last pushed today), licensed Apache-2.0. It adds 16 tokens to every session and 1,152 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
code-reviewer
当一个主要项目步骤完成并需要根据原始计划和编码标准进行审查时使用此智能体。示例: Context: 用户正在创建一个代码审查智能体,应在逻辑代码块编写完成后调用。user: "我已经按照计划第 3 步完成了用户认证系统的实现" assistant: "干得好!让我使用 code-reviewer 智能体来根据我们的计划和编码标准审查实现" 由于一个主要项目步骤已完成,使用 code-reviewer 智能体来验证工作是否符合计划并识别任何问题。 Context: 用户完成了一个重要功能的实现。user: "任务管理系统的 API 端点现在完成了——这涵盖了我们架构文档中的第 2 步" assistant: "很好!让我用…
code-reviewer
Senior Android code reviewer that evaluates changes across five dimensions — correctness, readability, architecture, security, performance. Use for thorough code review before merge.
Demonstrate
Agent for demonstrating VS Code features.
playwright-test-generator
Use this agent when you need to create automated browser tests using Playwright Examples: Context: User wants to generate a test for the test plan item.
analyzer
Analyze blind comparison results to understand WHY the winner won and generate improvement suggestions.
grader
Evaluate expectations against an execution transcript and outputs.