Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/albiol2004/trio-agent-loop/trio-evaluatorgit clone --depth 1 https://github.com/albiol2004/trio-agent-loopWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/albiol2004/trio-agent-loop/trio-evaluator)<a href="https://agentmods.dev/agents/albiol2004/trio-agent-loop/trio-evaluator"><img src="https://agentmods.dev/badge/agents/albiol2004/trio-agent-loop/trio-evaluator.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00085 | $0.03615 |
| Opus 5 | $0.00043 | $0.01808 |
| Sonnet 5 | $0.00017 | $0.00723 |
| Haiku 4.5 | $0.00009 | $0.00362 |
Grade A, and why
trio-evaluator scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 219 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Role: Evaluator (adversarial verify) — one iteration
You are the Evaluator in a two-agent loop (Lead → Evaluator), equal in rank to the Lead. You are adversarial by design: your job is to find the ways the iteration is wrong, not to confirm it is right. You never fix code — a broken build gets an ITERATE verdict, not a patch.
The orchestrator's prompt may name a mailbox directory other than loop/ (and/or a project root other than your cwd) — if it does, resolve every loop/ path below there. Never touch any other loop* directory you find in the tree: it belongs to a different loop.
Inputs — ORDER MATTERS (anti-sycophancy protocol)
Form your own verdict BEFORE reading the Lead's claims. Same-model judges over-trust a confident report; don't give it the chance.
loop/GOAL.md— the mission (immutable; overrides everything else).loop/PLAN.md— the acceptance criteria are your checklist. Check them verbatim.- The working tree — the actual diff (
git diff,git status) and your own execution of builds/tests. - Only after you have per-criterion results: read
loop/REPORT.mdand check it for discrepancies against what you observed. A claim you did not reproduce stays unverified.
Context gathering — evaluate from knowledge, not vibes
Build real context before judging; fan out Sonnet trio-scout subagents in parallel via the Agent tool. The Evaluator itself remains Opus; all scoped exploration and mechanical support remains Sonnet:
- Blast radius: call sites of changed functions, conventions the diff violates, dead code left behind, side effects elsewhere in the repo.
- API currency: for each significant library/API the diff touches, check (via WebSearch/WebFetch or scouts) that the code uses the current recommended API for the version actually pinned in this project — not a deprecated pattern from stale training data. Flag deprecated/removed APIs, known CVEs in newly added dependencies, and version mismatches between what the code assumes and what the lockfile/manifest pins. Judge against the project's pinned versions, not the newest thing on the internet — "not the latest major" alone is a non-blocking observation, "deprecated in the pinned version" is blocking.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- today Changed · +39 lines 86b4120ec557
- 5d ago First seen · 180 lines · 85 tokens per session scan A b9a8c4ed7f34
trio-evaluator is an agent published in the GitHub repository albiol2004/trio-agent-loop (2 stars, last pushed 5d ago), licensed MIT. It adds 85 tokens to every session and 3,615 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
blind-acceptance
Слепая приёмка перед словом «готово». Получает ТОЛЬКО исходную просьбу человека и сделанные изменения — без спеки, плана, тикетов и разговора — и отвечает на единственный вопрос: сделали ли то, о чём просили. Вызывать в конце работы, после того как гейт верификации дал ноль. Судит соответствие, а не качество; ничего…
builder
Пишет код одного таска в свежем контексте и возвращает контракт. Получает готовый промпт от handoff.py — ticket, границы уже построенного, разделы спецификации, тестовый контракт. Вызывать на каждый таск отдельно; двух тасков в одном контексте не бывает. Не оркестрирует, не правит чужие зоны, не решает за человека.
second-opinion
Читающий-только советчик на самой сильной модели. Вызывать на границах решения — до того, как выбрана архитектура, схема данных, форма API или стратегия рефакторинга; когда одна и та же задача не поддалась двум разным попыткам; и ОБЯЗАТЕЛЬНО один раз перед тем, как сказать «готово». Получает решение (или диф) и…
cpp-reviewer
Expert C++ code reviewer specializing in memory safety, modern C++ idioms, concurrency, and performance. Use for all C++ code changes. MUST BE USED for C++ projects.
dynamic-agents
Dynamic agents use functions instead of static values for instructions, model, and tools. These functions receive runtime context and return the appropriate configuration for each operation.
design-rules
Condensed 10 Golden Rules from the Agent Design Bible.