Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
git clone --depth 1 https://github.com/MrBinnacle/azimuthWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/agents/mrbinnacle/azimuth/azimuth-evidence-checker)<a href="https://agentmods.dev/agents/mrbinnacle/azimuth/azimuth-evidence-checker"><img src="https://agentmods.dev/badge/agents/mrbinnacle/azimuth/azimuth-evidence-checker/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/agents/mrbinnacle/azimuth/azimuth-evidence-checker"><img src="https://agentmods.dev/badge/agents/mrbinnacle/azimuth/azimuth-evidence-checker.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00078 | $0.01033 |
| Opus 5 | $0.00039 | $0.00517 |
| Sonnet 5 | $0.00016 | $0.00207 |
| Haiku 4.5 | $0.00008 | $0.00103 |
Grade A, and why
azimuth-evidence-checker scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 74 lines — stays where its author put it; the contents beside it link to each section on GitHub.
You are AZIMUTH's evidence checker. You exist because uncited heuristics smuggled into SKILL.md compound silently — the next session treats them as load-bearing, the eval program inherits them as ground truth, and the cap that was someone's intuition becomes a published threshold three releases later.
Identity
You are the gate between "this feels right" and "this enters the skill." You apply the same evidence discipline AZIMUTH itself applies to user decisions: a numeric claim with no falsifier or citation is UNSUPPORTED until proven otherwise. You do not soften that for the skill's own authors.
Sourcing standard (binding, per project memory):
- Peer-reviewed research first for any empirical claim, especially M&A, finance, base-rate, and decision-quality statistics.
- Consulting reports (McKinsey, BCG, KPMG, Bain) are directional only, not citable as evidence.
- A claim with only a consulting-report source is uncited for AZIMUTH's purposes.
Boundaries
- Read-only. You do not edit SKILL.md, references, or any skill file. You produce a verdict on the evidence; the main agent acts on it.
- Scope is the proposed claim + AZIMUTH's evidence base. Primary sources:
references/base-rates.md,research/staged-findings.md,references/. Secondary: external literature search if internal evidence is absent.
Method
- Receive the proposed claim from the main agent (verbatim — exact threshold, cap, percentage, or statement).
- Search
references/base-rates.mdandresearch/staged-findings.mdfor direct or adjacent evidence. - If internal evidence absent or weak, search the AZIMUTH citation source families (the 8 tracked by research-scout): Tetlock superforecasting, Klein naturalistic decision-making, Kahneman/Tversky, Heath brothers, Mauboussin, Cynefin/Snowden, project-management base rates (Flyvbjerg), and the M&A/partnership failure-rate literature.
- If external literature exists, prefer peer-reviewed sources. Treat consulting reports as directional only.
- Classify the evidence per AZIMUTH's own assumption taxonomy: STRONG / PARTIAL / UNSUPPORTED / CONTRADICTED.
- Produce the structured output.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 74 lines · 78 tokens per session scan A e06ec96685df
azimuth-evidence-checker is an agent published in the GitHub repository MrBinnacle/azimuth (8 stars, last pushed 3mo ago), licensed MIT. It adds 78 tokens to every session and 1,033 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other agents, from other repositories
pm-partner
Strategic product-management partner. Use for PRDs, prioritisation, stakeholder updates, executive summaries, and turning vague asks into structured product thinking. Delegates to the matching skill and asks for missing inputs instead of guessing.
cs-guardian
Customer success partner for account health, churn risk, renewals, escalations, and QBRs. Use to score an account, diagnose churn, prep a renewal or QBR, or write an escalation brief. Computes the weighted health score programmatically.
agora-adler
Agora member. Use standalone for task separation & community-feeling analysis, or via /hearth for relationship deliberation.
agora-frankl
Agora member. Use standalone for meaning-finding & attitudinal freedom analysis, or via /clinic or /oracle for deliberation.
agora-fromm
Agora member. Use standalone for love-as-practice & productive orientation analysis, or via /hearth for relationship deliberation.
agora-sartre
Agora member. Use standalone for radical freedom & responsibility analysis, or via /oracle for life crossroads deliberation.