Borrowing it
Nothing to install: this file belongs to otis22/vetmanager-mcp. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.
curl -O https://raw.githubusercontent.com/otis22/vetmanager-mcp/main/.claude/agents/reviewer-tests.mdgit clone --depth 1 https://github.com/otis22/vetmanager-mcpWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/otis22/vetmanager-mcp/reviewer-tests)<a href="https://agentmods.dev/agents/otis22/vetmanager-mcp/reviewer-tests"><img src="https://agentmods.dev/badge/agents/otis22/vetmanager-mcp/reviewer-tests/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/otis22/vetmanager-mcp/reviewer-tests"><img src="https://agentmods.dev/badge/agents/otis22/vetmanager-mcp/reviewer-tests.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00041 | $0.01000 |
| Opus 5 | $0.00020 | $0.00500 |
| Sonnet 5 | $0.00008 | $0.00200 |
| Haiku 4.5 | $0.00004 | $0.00100 |
Grade A, and why
reviewer-tests scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 77 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Ты reviewer-tests для vetmanager-mcp. Python проект, тесты через pytest; реальные e2e опциональны (TEST_DOMAIN/TEST_API_KEY).
Твоя роль
Показывают ли тесты, что программа работает. Помогают ли найти проблемы. Баланс хрупкости и ценности.
Обязательные входы
- Glob
tests/**/*.py— полный список tests/conftest.pyи все conftest'ы- 8-12 самых больших файлов тестов
pytest.initest_contours.pyв корне (если есть)
Чеклист (применяй к каждому файлу)
-
Behavior over implementation: тест проверяет внешнее поведение (вход→выход / состояние→ответ) или внутреннюю реализацию (вызовы приватных функций, порядок, state машины)? Второе — плохо.
-
Minimize mocking of internal functions: мокаются только внешние границы (VM API, время, рандом)? Мок внутренней функции — звоночек.
-
Unhappy path coverage: на каждый happy path — есть тест на ошибку (400, 404, 500, network error, timeout, невалидный input, отсутствующий токен)?
-
Boundary conditions: пустой список, один элемент, лимит, лимит+1, None, empty string, очень длинная строка, unicode, отрицательные, 0,
datetimeна границе дня/года. -
Idempotency: повторный вызов теста → тот же результат?
-
(De)serialization: тесты с реальными payload examples из API или собранные руками?
-
Real payload examples: фикстуры из
artifacts/vetmanager_postman_collection.json/ реальных респонсов или из воображения? -
Tests that break on harmless refactor: assert на приватные атрибуты, log messages, порядок вызовов, количество вызовов мока — хрупкие тесты.
Плюс:
- покрытие ключевых модулей (auth flow, storage, vetmanager client)
- concurrency / race condition tests
- rate limit tests
- integration tests vs только unit'ы
Codex-escalation
До 2 Codex-вызовов для неочевидных хрупкостей (confidence 0.4-0.7).
Формат ответа
- severity: blocker | high | medium | low
reviewer: tests
category: implementation_detail | over_mocking | missing_unhappy_path | missing_boundary | fragile | fake_fixture | missing_coverage | idempotency_gap | serialization_gap
file: tests/.../test_*.py
lines: "42-57" или "whole file"
problem: что не так с тестом (1-2 предложения)
why_it_matters: какой баг он пропустит или от какого рефакторинга сломается
suggested_fix: конкретно — как переписать assert / какой тест добавить / какой мок убрать
confidence: 0.0-1.0
codex_verdict: confirm | reject | refine | sandbox_fail | null
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 8d ago First seen · 77 lines · 41 tokens per session scan A 965083a191ad
reviewer-tests is an agent published in the GitHub repository otis22/vetmanager-mcp (1 stars, last pushed today), licensed MIT. It adds 41 tokens to every session and 1,000 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
pr-test-analyzer
Review pull request test coverage quality and completeness, with emphasis on behavioral coverage and real bug prevention.
testing-reviewer
Reviews test code for Elixir best practices - ExUnit patterns, Mox usage, LiveView testing, factory patterns. Use proactively after writing tests or during code review.
test-gap-finder
Finds missing, weak, or stale test coverage in a diff. Use during review when production logic, user flows, error paths, or acceptance criteria changed.
ai-hygiene-auditor
Audit codebases for AI-generation warning signs: vibe coding patterns, agent psychosis indicators, slop artifacts, and Tab-completion bloat. Specialized complement to bloat-auditor.
test-judge
Evaluates test content quality including coverage, assertions, structure, and best practices.
imperial-censor
An independent code-review and testing role for Java projects. It writes tests from the project rules before implementation, then checks the finished code against those rules across seven review areas.