Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add debabsah/analytics-office --skill audit-my-assumptionsgit clone --depth 1 https://github.com/debabsah/analytics-officeWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/debabsah/analytics-office/audit-my-assumptions)<a href="https://agentmods.dev/skills/debabsah/analytics-office/audit-my-assumptions"><img src="https://agentmods.dev/badge/skills/debabsah/analytics-office/audit-my-assumptions/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/debabsah/analytics-office/audit-my-assumptions"><img src="https://agentmods.dev/badge/skills/debabsah/analytics-office/audit-my-assumptions.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00249 | $0.03610 |
| Opus 5 | $0.00125 | $0.01805 |
| Sonnet 5 | $0.00050 | $0.00722 |
| Haiku 4.5 | $0.00025 | $0.00361 |
Grade A, and why
audit-my-assumptions scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 77 lines — stays where its author put it; the contents beside it link to each section on GitHub.
audit-my-assumptions
The colleague who makes you write down what you're taking for granted before you build a quarter of work on top of it — then tries to break the load-bearing ones while it's still cheap.
When to use
Fire FIRST — before you build/derive on it AND before you present/stand behind it. Fire when about to build or derive on top of something you did not author and fully verify, OR when about to present a number (of any provenance — inherited, self-derived, or long-trusted) and you want its premises checked before the room: rebuild a report from inherited stored procs, derive a new metric from an existing query, stack more years onto an old workbook, turn a proc's output into a deck, re-point a model at a new source. The question is "what am I silently assuming, and which of those assumptions would poison everything if it's wrong?" — asked before the build, not after a stakeholder squints at the output.
Do NOT fire when a number is already in hand and known wrong (that's triage-my-number, downstream), to pin what a metric should mean (kpi-contract), to review one code object against a contract (review-my-query), or to audit a whole accreted knowledge base (kb-reconcile). This runs at the front of the work, on the inputs.
This vs. its neighbors. The gate is no symptom yet: you're about to build/derive on, or present/stand behind, a number whose load-bearing premises are unexamined (any provenance — inherited, self-derived, or long-trusted) and you want them falsified first. Route elsewhere when the case differs — a number already known wrong → triage-my-number (diagnose the whole failure surface); a query's code logic suspect while the definition is trusted → review-my-query; your whole accreted knowledge base, not one source → kb-reconcile. Among the present-moment trio — premises → here · phrasing → brief-my-findings · pushback → defend-my-number — fire here for "what am I assuming under this number?" (it can be clean and still wrong); for "how do I write this up?" use brief-my-findings; for "how do I hold up under live challenge?" use defend-my-number.
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 77 lines · 0 tokens per session scan A 68ad4f2c6d80
audit-my-assumptions is a skill published in the GitHub repository debabsah/analytics-office (9 stars, last pushed 3mo ago), licensed MIT. It adds 249 tokens to every session and 3,610 once invoked, about $0.0012 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
brooks-sweep
Full-sweep mode: runs a unified analysis across all quality dimensions — code decay, architecture, tech debt, and test quality — then applies fixes directly to the codebase. Safe changes are auto-applied; risky changes are confirmed before execution. Drawing on twelve classic engineering books. Triggers when: user…
critical-code-reviewer
Rigorously review code or pull requests for correctness, security, accessibility, maintainability, tests, and edge cases. Use when users request a critical code review, want a guided walkthrough of findings, need implementer-facing feedback, or want to prepare, create, or submit a GitHub pull request review.
second-pass-review
Independent audit of sanitized specs in workspace/output/. Three parallel LLM-based reviewer roles check structural leakage, content contamination, and behavioral completeness. Run AFTER Layer 5 sanitization, BEFORE implementation handoff.
check-pr
Read-only inspection of a single GitHub PR lifecycle — checks CI, review threads, description sync, and mergeability, and returns PASS or FAIL with per-gate findings. Never invokes the merge button. Use when verifying a PR is ready to merge, polling lifecycle progress, checking mergeability, or babysitting a GitHub PR…
codex-review
Get an independent code review from OpenAI Codex (GPT-5) on uncommitted changes, a branch diff, a PR, or a specific module — returns findings grouped by severity with file:line references. Use when the user says "review this", "review my changes", "check this PR", "what did I miss", or before they commit or merge…
github-pr-creation
Creates GitHub Pull Requests with automated validation and task tracking. Use when user wants to create PR, open pull request, submit for review, or check if ready for PR. Analyzes commits, validates task completion, generates Conventional Commits title and description, suggests labels. NOTE - for merging existing…