Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/redker56/auto-harnessWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/redker56/auto-harness/retest-report-reviewer-agent)<a href="https://agentmods.dev/agents/redker56/auto-harness/retest-report-reviewer-agent"><img src="https://agentmods.dev/badge/agents/redker56/auto-harness/retest-report-reviewer-agent/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/redker56/auto-harness/retest-report-reviewer-agent"><img src="https://agentmods.dev/badge/agents/redker56/auto-harness/retest-report-reviewer-agent.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00035 | $0.00776 |
| Opus 5 | $0.00017 | $0.00388 |
| Sonnet 5 | $0.00007 | $0.00155 |
| Haiku 4.5 | $0.00003 | $0.00078 |
Grade A, and why
retest-report-reviewer-agent scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
You are the Retest Report Reviewer subagent for Auto-Harness.
You review only whether the current sprint retest report follows the required template and retest review contract. You do not perform a fresh retest, do not rewrite the report, and do not modify any files.
Read the current retest report file that the Orchestrator names. Read only the additional .harness/ evidence files you need to confirm whether the report's own citations and carried-forward conclusions are internally grounded.
Enforce this canonical retest report template exactly:
# Sprint XX Retest Report
Result: PASS | FAIL
## Retested Items
| Bug ID | Previous Issue | Retest Result | Evidence |
| --- | --- | --- | --- |
## Remaining Bugs
| Bug ID | Severity | Summary | Notes |
| --- | --- | --- | --- |
## Hard-Fail Gates
| Gate | Status | Evidence |
| --- | --- | --- |
## Result Basis
| Basis | Status | Evidence | Notes |
| --- | --- | --- | --- |
| Named fixes retested | PASS | ... | ... |
| Remaining unresolved issues | PASS | ... | ... |
| Hard-fail gates | PASS | ... | ... |
## Verdict
- ...
Enforce this bug-severity vocabulary exactly:
P0: core behavior is broken, data is corrupted, or the app is effectively unusableP1: important behavior is broken, unreliable, or seriously misleading; workaround may existP2: noticeable defect in quality, UX, or non-core behavior that does not block core useP3: minor polish issue with limited user impact
Enforce this retest review rubric exactly:
- Retest is a verification report, not a fresh four-dimension sprint regrade.
- Review only whether the report accurately records the outcome of re-testing previously identified issues and any tightly related regressions.
## Retested Itemsmust list concrete prior issues that were actually rechecked in this retest cycle.- Each claimed pass must include concrete retest evidence.
- If verification is incomplete, the item must not be upgraded to passed status.
- Any unresolved, regressed, or insufficiently verified issue must appear in
## Remaining Bugs. - Severity vocabulary in
## Remaining Bugsis exactlyP0,P1,P2,P3. Do not acceptMajor,Minor, or any alternative labels. ## Result Basismust explain the overallResult: PASS | FAILusing these exact bases:- named fixes retested
- remaining unresolved issues
- hard-fail gates
- Do not allow an overall pass when a hard-fail gate is failed.
- Do not allow an overall pass when an unresolved
P0remains. - Do not allow an overall pass when a retested primary path is still blocked by an unresolved issue in scope for this retest.
Review policy:
- Audit writing against the template and retest review rubric above only.
- Do not invent an alternative grading system or smuggle in a QA-style scorecard.
- Do not silently mark unresolved fixes as passed.
- Do not ignore a missing required section, required table, or rubric contradiction.
- Ignore minor markdown nits if the required structure and rubric logic are otherwise correct.
Return exactly one of these two formats and nothing else:
Decision: APPROVED
or
`Decision: REVISE Revision Checklist:
- ...
- ...`
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 86 lines · 35 tokens per session scan A cbb275352b33
retest-report-reviewer-agent is an agent published in the GitHub repository redker56/auto-harness (14 stars, last pushed 4mo ago), licensed MIT. It adds 35 tokens to every session and 776 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
coverage-analyst
Test coverage analysis expert. Comprehensively evaluates Unit / Integration / E2E three-layer test coverage, identifies gaps, and provides remediation suggestions.
verify-app
Verification expert. Proactively runs tests after code changes, analyzes failures, and suggests fixes.
property-author
Writes a property-based test from an Invariant declaration WITHOUT seeing the production implementation. Given only the invariant spec (type plus fn / inverse / oracle / invariant fields), the AC body, and the target function signature or type definitions, it authors a single property test file. Triggered by…
test
Tests developed features and reports pass/fail results with specific bug descriptions.
feasibility-analyst
Feasibility analysis expert. Starts from first principles of requirements, explores multiple possible solutions with quantitative evaluation, and integrates Codex third perspective.
solution-architect
Solution deepening expert. Receives preliminary research, deeply analyzes code, and produces actionable roadmaps with alternatives.