Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/uditgoenka/autoresearch/debuggit clone --depth 1 https://github.com/uditgoenka/autoresearchWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00021 | $0.01022 |
| Opus 5 | $0.00010 | $0.00511 |
| Sonnet 5 | $0.00004 | $0.00204 |
| Haiku 4.5 | $0.00002 | $0.00102 |
Grade A, and why
autoresearch:debug scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
Copies of this mod
1 near-identical copy found in the catalogue:
- autoresearch_debug — 97% identical, 4 lines differ
How it starts
The opening of the file, as written. The whole thing — 98 lines — stays where its author put it; the contents beside it link to each section on GitHub.
EXECUTE IMMEDIATELY.
Parse Arguments
Extract from $ARGUMENTS:
Scope:or--scope— file globs to investigateSymptom:or--symptom— error message or behavior descriptionIterations:or--iterations— default 15. "unlimited" for unbounded.--fix— shorthand for--chain fix--severity— filter: critical, high, medium, low--technique— force specific technique--evals,--evals-interval N,--chain
Setup (if required context missing)
If Scope and Symptom both missing:
- Auto-scan: run tests, lint, typecheck to detect existing failures
- AskUserQuestion (single batch): Q1 (Issue): "What's the problem?" — hunt all bugs, specific error, failing tests, CI failure, performance Q2 (Scope): "Which files?" — suggested globs + entire codebase Q3 (Depth): "How deep?" — quick (5), standard (15), deep (30+), unlimited Q4 (After): "When bugs found?" — report only, find and fix (--chain fix), chain to other, ask each time If all provided → skip.
Investigation Techniques
| Technique | When to Use |
|---|---|
| Binary search | Know when it worked, find when it broke |
| Differential | Compare working vs broken state |
| Minimal reproduction | Simplify to smallest failing case |
| Trace | Follow execution path through code |
| Pattern search | Grep for known anti-patterns |
| Working backwards | Start from error, trace to root cause |
Establish Baseline (before loop)
- Auto-scan for failures if no symptom provided
- Create output directory:
autoresearch/debug-{YYMMDD}-{HHMM}/ - TSV header:
# metric_direction: higher_is_better\niteration\ttimestamp\thypothesis\tstatus\ttechnique\tevidence\tfile_line - Metric = cumulative confirmed findings count
Iteration Loop
Phase 1: Review Context
- Read results TSV (past findings)
- Assess: what's been tested, what vectors remain
- If no hypotheses left → early stop
Phase 2: Hypothesize
- Form ONE specific, falsifiable hypothesis
- Format: "I hypothesize that {X} because {evidence}. Test by {Y}."
- Hypothesis must be testable and different from all previous
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 98 lines · 21 tokens per session scan A 61fe3ef78390
autoresearch:debug is a command published in the GitHub repository uditgoenka/autoresearch (5,966 stars, last pushed 19d ago), licensed MIT. It adds 21 tokens to every session and 1,022 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
drift
Read the Genesis build phases document: docs/architecture/genesis-v3-build-phases.md.
challenge
Read the specified design doc section.
thoth:dashboard
Alias for status --dashboard; manage the local dashboard backed by .thoth ledgers.
thoth:doctor
Alias for status --doctor; strictly audit project health without writing authority.
feature
End-to-end feature/bug-sweep workflow for aitm — understand, reproduce against a real run, explore and build with a hive of parallel agents in this one checkout (never worktrees), path-disjoint slices, verify under Bun AND Node, PR, merge, and (only when asked) release to npm. Tracks in GitHub issues. Reads intent…
merge-pr
Drive an open PR to merge, then advance to the next PR group.