Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add CassioRoos/godfly-skills --skill competing-hypothesesgit clone --depth 1 https://github.com/CassioRoos/godfly-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/cassioroos/godfly-skills/competing-hypotheses)<a href="https://agentmods.dev/skills/cassioroos/godfly-skills/competing-hypotheses"><img src="https://agentmods.dev/badge/skills/cassioroos/godfly-skills/competing-hypotheses/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/cassioroos/godfly-skills/competing-hypotheses"><img src="https://agentmods.dev/badge/skills/cassioroos/godfly-skills/competing-hypotheses.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00114 | $0.00999 |
| Opus 5 | $0.00057 | $0.00500 |
| Sonnet 5 | $0.00023 | $0.00200 |
| Haiku 4.5 | $0.00011 | $0.00100 |
Grade A, and why
competing-hypotheses scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured today.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 115 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Analysis of Competing Hypotheses
Don't pick the first plausible approach. Map the full solution space. Evaluate each against evidence. Find the one that best explains all the constraints.
Core Principle
The CIA developed ACH because analysts kept falling for confirmation bias -- finding evidence that supported their first hypothesis and ignoring the rest. The same happens in engineering. You think of an approach, find reasons it works, and stop looking. ACH forces you to consider ALL plausible approaches and evaluate them against ALL evidence.
When to Use
- Choosing between architectures, frameworks, or libraries
- When "the obvious choice" hasn't been compared to alternatives
- When team members disagree on approach
- When the user needs counterpoints to their proposed solution
- When a technology bet has limited reversibility
The Process
Step 1: List All Plausible Hypotheses
Compare the genuinely credible approaches; 3-5 is a useful default, not a minimum. Two may be enough. Do not invent alternatives to fill a quota.
Rules:
- Each must be technically credible; cite real use when available and label unproven approaches rather than inventing production precedent.
- Include the user's proposed approach
- Look for a materially different option the user may not have considered; include it only if credible.
- Consider "do nothing" or "simplest possible"; explain if constraints rule it out.
Step 2: List All Evidence
Gather evidence from all available sources:
- Requirements and constraints
- Codebase patterns and existing architecture
- Team capabilities and timeline
- Performance requirements and scale targets
- Real-world usage data from similar systems
- Known failure modes of each approach
Step 3: Build the Diagnostic Matrix
| Evidence | Approach A | Approach B | Approach C |
|---|---|---|---|
| [constraint 1] | ++ | - | + |
| [constraint 2] | - | ++ | + |
| [constraint 3] | + | + | -- |
Ratings: ++ strongly supports, + supports, o neutral, - contradicts, -- strongly contradicts
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- today Changed · -1 lines f2f0b7628af6
- yesterday Changed · -5 lines · +6 tokens per session 82021ca751cb
- 10d ago First seen · 121 lines · 108 tokens per session scan A d6cd01ac6f39
competing-hypotheses is a skill published in the GitHub repository CassioRoos/godfly-skills (1 stars, last pushed today), licensed MIT. It adds 114 tokens to every session and 999 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
auto-test-code
A structured process for critically reviewing and testing software code. It records review findings, test plans, commands, results, and supporting files in a project workspace.
git-pr-review
A read-only reviewer for GitHub pull requests, which are proposed code changes submitted for review. It produces an evidence-based report about whether a pull request should be merged.
code-remediate
Apply selected review fixes; bare PR targets use current online items, while PR +review adds the latest matching artifact.
code-review
Close PRs at an evidence gate or review local diffs/PRs with specialists and JSON artifacts.
code-reviewer
A code-review workflow that checks completed work against its requirements or plan before merging. It groups findings by severity and gives each issue a fix and a way to verify it.
coding-assistant
Provides coding assistance with best practices and code review.