Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/pedromosquera/squadai/systematic-debuggingnpx skills add PedroMosquera/squadai --skill systematic-debugginggit clone --depth 1 https://github.com/PedroMosquera/squadaiWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00015 | $0.00669 |
| Opus 5 | $0.00008 | $0.00334 |
| Sonnet 5 | $0.00003 | $0.00134 |
| Haiku 4.5 | $0.00002 | $0.00067 |
Grade A, and why
systematic-debugging scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 96 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Systematic Debugging Skill
Diagnose and fix failing tests using a structured 4-phase protocol. This skill is used by the Debugger when tests are failing and the cause is not immediately obvious.
The 4-Phase Protocol
Phase 1: REPRODUCE
-
Identify the failing test(s) exactly.
- Run the test suite and capture the exact failure output
- Note the assertion that failed and the expected vs actual values
- Record whether the failure is consistent or flaky
-
Create a minimal reproduction.
- Isolate the failing test from its suite
- Remove any unrelated setup or fixtures
- Confirm the minimal test still fails
-
Document what you know.
- When did the test start failing? (after which commit?)
- Does it fail on all machines or just some environments?
- Is the failure deterministic?
Phase 2: ISOLATE
-
Identify the failure boundary.
- Is the bug in the test itself (wrong assertion) or in the code under test?
- Add temporary logging or debug output at key points
- Use a binary search approach: remove half the code, check if still failing
-
Check common causes first.
- Off-by-one errors in loops
- Nil pointer dereference or missing initialization
- Incorrect error handling (error swallowed, wrong type)
- State leaking between tests (missing cleanup)
- Race condition (run with
-raceflag) - Wrong mock or stub behavior
-
Trace the data flow.
- Follow the input through each transformation
- Compare the actual intermediate values against expected values
- Find the first point where actual diverges from expected
Phase 3: FIX
-
Make the smallest targeted fix.
- Change ONLY what is necessary to fix the identified root cause
- Do not refactor or clean up while fixing
- If the fix requires a larger change, note it but keep it separate
-
Verify the fix addresses the root cause.
- Explain in a comment why the fix works
- Ensure the fix does not introduce new edge cases
-
Run the previously-failing test — it must now pass.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 96 lines · 15 tokens per session scan A d490a9584b8d
systematic-debugging is a skill published in the GitHub repository PedroMosquera/squadai (8 stars, last pushed 1mo ago), licensed MIT. It adds 15 tokens to every session and 669 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
etsy-category-listing
Etsy category page scraper: given an Etsy category URL (e.g. https://www.etsy.com/c/jewelry) and optional page number, returns paginated product listings with listingId, shopId, title, url, image, salePrice, originalPrice, currency, rating, reviewCount, shopName, isAd, freeShipping, badge from category and subcategory…
human-approval
Request human approval before performing a SAFETY-CRITICAL, IRREVERSIBLE, or SCOPE-EXPANDING action — submit a structured context (action, scope, risk, consequence) plus options, then STOP the current turn. The platform redispatches the agent after the human decides. NEVER use for routine deliverables (writing docs /…
firebase-analytics
Use when logging analytics events, setting user properties, configuring default event parameters, building funnels, or adding screen-view tracking.
firebase-remote-config
Use when implementing feature flags, running A/B tests, setting parameter defaults, fetching/activating config, or enabling real-time config updates.
git-master
MUST USE whenever a task needs a commit or git-history investigation. Covers atomic commits, staging, commit-message style, rebase, squash, fixup/autosquash, blame, bisect, reflog, git log -S/-G, and questions like who wrote this or when was this added. Do not use for ordinary code edits unless the user asks for git…
prismer-evolve-record
Record the outcome of applying an evolution strategy. Use after resolving an error where prismer-evolve-analyze provided a recommendation, to feed back success or failure to the network.