Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Rkamirage/consulting-research-to-output --skill build-consulting-evidence-basegit clone --depth 1 https://github.com/Rkamirage/consulting-research-to-outputWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/rkamirage/consulting-research-to-output/build-consulting-evidence-base)<a href="https://agentmods.dev/skills/rkamirage/consulting-research-to-output/build-consulting-evidence-base"><img src="https://agentmods.dev/badge/skills/rkamirage/consulting-research-to-output/build-consulting-evidence-base/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/rkamirage/consulting-research-to-output/build-consulting-evidence-base"><img src="https://agentmods.dev/badge/skills/rkamirage/consulting-research-to-output/build-consulting-evidence-base.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00098 | $0.01492 |
| Opus 5 | $0.00049 | $0.00746 |
| Sonnet 5 | $0.00020 | $0.00298 |
| Haiku 4.5 | $0.00010 | $0.00149 |
Grade A, and why
build-consulting-evidence-base scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 80 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Build Consulting Evidence Base
Test one bounded question. Preserve the central lead's hypothesis, decision relevance, evidence need, route, support/refute/inconclusive conditions, and return contract. Do not rebuild the issue tree or decide the answer.
When invoked directly, ask only for a missing access, permission, recipient, or research-boundary fact that changes the work; otherwise state an assumption and proceed.
Match the evidence route to the decision need
Do not ask the user or central lead to choose a research level. Start from the bounded hypothesis and required knowledge grain. Expand original-source pursuit, rival testing, direct observation, or governed records only when a high-impact claim lacks direct evidence, a material conflict remains, the decision is hard to reverse, or one more feasible route could change the answer.
For each leaf, predeclare:
decision criterion → diagnostic question → falsifiable leaf test
Name its knowledge kind—claim verification, mechanism understanding, buyer/workflow understanding, economics, or implementation reality—and required grain: instance, aggregate, category, or derived. Category evidence cannot by itself prove an instance workflow, mechanism, economics, performance, or implementation reality.
Research by hypothesis
- Map the evidence landscape, upstream sources, vocabulary, conflicts, and critical gaps briefly.
- Open originals suited to the knowledge kind: primary documents/data, direct user or buyer evidence, implementation records, product trials, transactions, or authoritative regulation. Snippets and AI summaries are discovery only.
- Extract the exact fact, population, period, unit, denominator, method, qualifier, applicability, and concrete instance or derived result required by the test.
- Seek the strongest rival, counterexample, contradiction, and failure condition.
- Run one targeted drill-down most likely to change the result or answer boundary.
- Stop when another focused source is unlikely to change the test result or its decision-relevant confidence.
What ships with it
9 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- agents/openai.yaml 295 B
- assets/evidence-packet.md 4.9 KB
- assets/governed-evidence-packet.md 5.3 KB
- assets/rapid-scan-packet.md 1.9 KB
- references/evidence-contracts.md 15 KB
- references/quality-bar.md 8.0 KB
- references/research-modes-and-tools.md 6.4 KB
- references/research-workflow.md 16 KB
- references/source-discipline.md 7.7 KB
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 12d ago First seen · 80 lines · 98 tokens per session scan A 2169a104a386
build-consulting-evidence-base is a skill published in the GitHub repository Rkamirage/consulting-research-to-output (4 stars, last pushed 1mo ago), licensed MIT. It adds 98 tokens to every session and 1,492 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
html-ppt
For consulting delivery work: turn diagnosis, frameworks, and project work into a client-adoptable action plan. Built around the core query "consulting-final-deck", with engagement manager judgment, buyer-ready proof, and this outcome: accept the recommendation and commit owners to the roadmap.
html-ppt-zhangzara-soft-editorial
A digital-transformation roadmap for a legacy insurer — the diagnosis, the sequenced bets, and the operating rhythm to land them. Built as a decision-grade consulting deck for client executives.
html-ppt-zhangzara-mat
A margin-recovery diagnosis for a regional grocery chain — the governing thought, the driver tree, the priorities, and the roadmap. Built as a decision-grade consulting deck for client sponsor, steering committee.
huashu-golden-circle
A brand-repositioning strategy for a heritage coffee chain — why, how, what — the governing idea and the moves to prove it. Built as a decision-grade consulting deck for client CMO, board.
huashu-luxe-whitespace
A market-entry study for a luxury skincare brand entering Asia — segmentation, positioning, channel, and the phased plan. Built as a decision-grade consulting deck for client leadership.
ppt-keynote
An operating-model redesign for a scaling logistics firm — the diagnosis, the target model, and the transition roadmap. Built as a decision-grade consulting deck for client sponsor, ops leaders.