debug-systematically

A structured debugging guide for investigating software errors and unexpected behaviour. It covers reproducing the problem, narrowing it to a component, and forming testable explanations.

In plain words
What is it for?
Use it when a command, test, or application fails. It helps document reproduction steps, test smaller parts, inspect recent changes, and check edge cases.
Why use it?
It reduces guesswork by making you confirm the failure and isolate its cause before changing code.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/eveld/claude/debug-systematically
Any agent
npx skills add eveld/claude --skill debug-systematically
Clone the repo
git clone --depth 1 https://github.com/eveld/claude

Made for: Claude Code, Codex.

Per session 22 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,123 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00022 $0.01123
Opus 5 $0.00011 $0.00562
Sonnet 5 $0.00004 $0.00225
Haiku 4.5 $0.00002 $0.00112

Measured yesterday against content hash 8c960716ef57, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

debug-systematically scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/debug-systematically/SKILL.md · 204 lines

How it starts

The opening of the file, as written. The whole thing — 204 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Debug Systematically

Follow a structured debugging process when encountering errors or unexpected behavior.

Debugging Process

1. Reproduce

Goal: Confirm the error consistently occurs

  • Run the failing command/test again
  • Document exact steps to reproduce
  • Note any error messages verbatim
  • Check if error is consistent or intermittent

Example:

# Reproduce the error
make test

# Document output
Error: undefined method 'foo' on line 42

2. Isolate

Goal: Narrow down to specific component/function

  • Identify which component is failing
  • Remove unrelated code to isolate issue
  • Check if error occurs with minimal input
  • Use binary search to find breaking change

Example:

# Test individual components
go test ./pkg/auth/handler_test.go -v -run TestLogin

# Isolate to specific function
go test ./pkg/auth -v -run TestLogin/valid_credentials

3. Hypothesize

Goal: Form testable theories about the cause

  • Based on error message, what could cause this?
  • What changed recently that might affect this?
  • What assumptions might be wrong?
  • What edge cases aren't handled?

Examples:

  • Hypothesis 1: Nil pointer - missing initialization
  • Hypothesis 2: Type mismatch - wrong function signature
  • Hypothesis 3: Race condition - concurrent access

4. Test Hypotheses

Goal: Verify each hypothesis systematically

  • Test one hypothesis at a time
  • Add logging/debugging to verify assumptions
  • Check related code for similar patterns
  • Look at test failures for clues

Example:

// Test hypothesis 1: nil pointer
if handler.service == nil {
    log.Printf("DEBUG: service is nil")
}

// Test hypothesis 2: check types
log.Printf("DEBUG: user type=%T, expected=*User", user)

5. Fix

Goal: Apply fix and verify it resolves the issue

  • Implement the fix for confirmed root cause
  • Run the originally failing test/command
  • Verify fix doesn't break other functionality
  • Clean up any debug logging

Example:

// Fix: Initialize service before use
func NewHandler() *Handler {
    return &Handler{
        service: NewAuthService(), // was missing
    }
}

Read the full file on GitHub · 204 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 204 lines · 22 tokens per session scan A 8c960716ef57

Subscribe to this mod's changes

debug-systematically is a skill published in the GitHub repository eveld/claude (10 stars, last pushed 6mo ago), licensed MIT. It adds 22 tokens to every session and 1,123 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

code-indexing-pipeline

How Infigraph turns source into a graph — adding a language (tree-sitter vs ANTLR grammar-plugin), cross-file call resolution, SCIP compiler-grade enrichment, and file-watch/reindex triage. Use when adding language support, debugging unresolved calls or SCIP import, or triaging stale index/watcher issues.

intuit/infigraph · 70 tokens

review-pr-against-issue

Review one or more PRs against the GitHub issue(s) they claim to fix, including fetching PR branches directly when gh can't reach github.com (e.g. gh is authenticated to an enterprise host instead). Use whenever asked "does this PR fix issue.

intuit/infigraph · 61 tokens

analysis-subsystems

How Infigraph's multi-repo/group mode and taint analysis work internally — HTTP contract extraction heuristics, cross-service edge linking, combined-graph merge, remote mode, plus taint's line-based tracking and sanitizer heuristic. Use when working on crates/infigraph-core/src/multi/ or src/taint/, or investigating…

intuit/infigraph · 80 tokens

opencli-sitemap-author

Use when creating or maintaining OpenCLI site sitemaps: agent-facing navigation, page-state, action, workflow, API-reference, pitfall, and fallback knowledge for a website. Use after browser exploration discovers durable site context, when a sitemap is stale, or when promoting local site knowledge into the repo.

jackwener/OpenCLI · 67 tokens

golden-rss

Use when testing the rss golden build.

yusufkaraaslan/Skill_Seekers · 12 tokens

omh-code-review

This is a Hermes-native code-review workflow skill.

rlaope/oh-my-hermes · 46 tokens