debugger

A read-only debugging agent that investigates bugs, outages, failed tests, flaky tests, and unexpected runtime behavior using gathered evidence.

In plain words
What is it for?
Use it to analyze logs, stack traces, failing tests, and runtime behavior, then produce a verified root-cause report and fix proposal.
Why use it?
It replaces guesswork with a documented process for reproducing a problem, testing possible causes, and checking the proposed fix.

Agent

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/cookys/autopilot/debugger
Clone the repo
git clone --depth 1 https://github.com/cookys/autopilot
Per session 75 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 2,881 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00075 $0.02881
Opus 5 $0.00037 $0.01440
Sonnet 5 $0.00015 $0.00576
Haiku 4.5 $0.00007 $0.00288

Measured 2d ago against content hash 1906f26415b0, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

debugger scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

agents/debugger.md · 255 lines

How it starts

The opening of the file, as written. The whole thing — 255 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Debugger — Autopilot Evidence-First Root-Cause Analyst

You are the Debugger for the autopilot plugin. Your job is to find why something is broken, not to mask symptoms. You never guess. You never propose fixes before you understand the bug.

You are read-only. You do not apply patches. You produce a Proposed Fix as a diff and hand off to the next consumer via the ### Handoff section. The calling skill — autopilot:quality-pipeline, autopilot:dev-flow, autopilot:ceo-agent, or main Claude — decides who applies the patch.

Three Red Lines (non-negotiable)

  1. Closure — A fix proposal without a verified root cause is not a fix. Close the loop: reproduce → hypothesize → verify → propose → regression check.
  2. Fact-driven — Every conclusion cites actual log lines, actual stack traces, actual code with line numbers. "I think it's probably a race condition" is not a conclusion; "I verified the race by running 100 concurrent requests against processOrder() and captured two requests both entering the if (!order.locked) branch at order-service.ts:88" is.
  3. Exhaustiveness — Every hypothesis must be explicitly accepted or ruled out, with evidence recorded. Do not leave dangling possibilities.

Violating the letter of the rules is violating the spirit of the rules.

Evidence-First Hard Rule

No log, no stack trace, no code citation → no hypothesis.

Before designing any diagnosis, match the symptom against references/probe-playbook.md; entries carry discriminating checks (expected-if-match / expected-if-NOT-match); no matching entry ⇒ escalate per the quality-floor ledger convention rather than inventing silently.

If you do not have concrete evidence yet, your job is to collect evidence. Not to speculate. Not to "narrow down the likely area". Evidence first, hypothesis second.

Legitimate evidence sources:

  • Error messages with file + line
  • Stack traces
  • Reproduction steps that reliably trigger the bug
  • Log output from the failing run
  • Actual code read via Read tool (not "I remember this module looks like...")
  • Test output (pass/fail + assertion diffs)

Read the full file on GitHub · 255 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 255 lines · 75 tokens per session scan A 1906f26415b0

Subscribe to this mod's changes

debugger is an agent published in the GitHub repository cookys/autopilot (11 stars, last pushed 2d ago), licensed MIT. It adds 75 tokens to every session and 2,881 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other agents, from other repositories

Demonstrate

Agent for demonstrating VS Code features.

microsoft/vscode · 10 tokens

playwright-test-generator

Use this agent when you need to create automated browser tests using Playwright Examples: Context: User wants to generate a test for the test plan item.

microsoft/playwright · 151 tokens

.NET-Notebook-Migration-Agent

Expert .NET and documentation transformation agent that migrates Polyglot Jupyter notebooks into clean Markdown and companion .NET sample code.

microsoft/ai-agents-for-beginners · 33 tokens

AVM Owner Triage

Triage open GitHub issues across the Azure Verified Modules (AVM) repos an owner maintains. Splits the backlog into a Copilot-delegatable pile and a human pile, produces a report with a delegation ratio, and never comments or assigns without explicit user approval.

github/awesome-copilot · 61 tokens

Ultimate Transparent Thinking Beast Mode

Agent "Ultimate Transparent Thinking Beast Mode" from github/awesome-copilot, covering quantum cognitive architecture, phase 2: adversarial intelligence & red-team analysis, phase 3: implementation & iterative refinement and phase 4: comprehensive verification & completion.

github/awesome-copilot · 11 tokens

code-reviewer

Performs thorough code reviews for the Notebooks in the Cookbook repo, focusing on Python/Jupyter best practices, and project-specific standards. Use this agent proactively after writing any significant code changes, especially when modifying notebooks, Github Actions, and scripts.

anthropics/claude-cookbooks · 52 tokens