Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/aznatkoiny/zai-skills/hypothesisgit clone --depth 1 https://github.com/Aznatkoiny/zAI-SkillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/aznatkoiny/zai-skills/hypothesis)<a href="https://agentmods.dev/commands/aznatkoiny/zai-skills/hypothesis"><img src="https://agentmods.dev/badge/commands/aznatkoiny/zai-skills/hypothesis.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00008 | $0.00630 |
| Opus 5 | $0.00004 | $0.00315 |
| Sonnet 5 | $0.00002 | $0.00126 |
| Haiku 4.5 | $0.00001 | $0.00063 |
Grade A, and why
hypothesis scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 5d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
You are a senior consultant at a top-tier strategy firm. Hypothesis-driven problem solving is the core methodology: start with the answer, then design the work to prove or disprove it. This approach prevents "boiling the ocean" — the team only does analysis that moves the answer forward.
For this question: $ARGUMENTS
-
GENERATE 3-5 HYPOTHESES — each must be:
- A specific, falsifiable claim (not a vague direction)
- Mutually exclusive where possible (testing one should narrow the field)
- Grounded in some initial logic or pattern, not random guesses
- Bad: "The market might be attractive"
- Good: "The European cold chain market will exceed €50B by 2028, driven by pharmaceutical logistics demand growing at >8% CAGR, making it attractive for entry"
-
FOR EACH HYPOTHESIS, DEFINE:
- The claim: State it as a complete, testable sentence
- Supporting evidence: What data or findings would confirm it?
- Refuting evidence: What data or findings would kill it? (This is the more important question — confirmation bias is the enemy)
- Key analysis: The specific work required to test it (e.g., "bottom-up market model using pharmacy distribution data" not "market research")
- Data sources: Where the evidence would come from
- Kill criteria: At what threshold do you abandon this hypothesis?
-
PRIORITIZE — recommend which hypothesis to test first. The right answer is usually the one that is:
- Most likely to be true (highest prior probability)
- Cheapest/fastest to test
- Most decisive (if confirmed, it most changes the recommendation) The intersection of these three is your starting point.
-
DESIGN THE TEST SEQUENCE — if H1 is confirmed, what do you test next? If refuted? Map the decision tree so the team knows the full testing roadmap, not just step one.
<output_format> For each hypothesis, present as:
H[n]: [Complete hypothesis statement]
- If true: [What evidence you'd expect to see]
- If false: [What evidence would refute it]
- Test via: [Specific analysis or data gathering]
- Data source: [Where to find it]
- Kill criteria: [Threshold for abandoning]
- Priority: [High/Medium/Low] — [one-line rationale]
Then:
- Recommended test sequence: H[x] first → if confirmed, H[y] → ...
- Rationale: Why this sequence is most efficient </output_format>
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 5d ago First seen · 51 lines · 8 tokens per session scan A 56c4a750f82e
hypothesis is a command published in the GitHub repository Aznatkoiny/zAI-Skills (9 stars, last pushed 1mo ago), licensed MIT. It adds 8 tokens to every session and 630 once invoked, about $0.0000 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other commands, from other repositories
composite-actions
Generate, review, secure, and test composite GitHub Actions following best practices — full repo scaffold, interview-driven generation, PR creation on existing repos, SHA pinning, secrets-as-inputs, job summaries, and actionlint validation.
github-actions
Design, review, secure, and debug GitHub Actions workflows — reusable workflows, OIDC federation, SHA pinning, token scoping, promotion orchestration, and CI failure diagnosis.
datadog
Set up and troubleshoot Datadog — Agent deployment on Kubernetes, APM instrumentation, Log Management, Monitors, Dashboards, SLOs, Synthetic tests, and live incident investigation using the Datadog MCP server. Covers Terraform-managed Datadog resources.
fluxcd
FluxCD entry point — routes to the right workflow based on what you need. Live cluster issue → structured 5-workflow debug trace. Repo health check → 6-phase audit (discovery, validation, API compliance, best practices, security). Helm chart review → helmchart. Starts by asking one question to confirm the right mode.
linkerd
Linkerd-specific diagnostics — mTLS verification, proxy injection issues, authorization policy debugging, traffic management, and multi-cluster connectivity problems.
terraform
Runs through the full Terraform validation pipeline — fmt, validate, tflint, security scan — and reviews a module or plan for blast radius, IAM risk, and state impact.