Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add agents/mubaidr/gem-team/gem-debuggergit clone --depth 1 https://github.com/mubaidr/gem-teamWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00021 | $0.00904 |
| Opus 5 | $0.00010 | $0.00452 |
| Sonnet 5 | $0.00004 | $0.00181 |
| Haiku 4.5 | $0.00002 | $0.00090 |
Grade A, and why
gem-debugger scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 114 lines — stays where its author put it; the contents beside it link to each section on GitHub.
DEBUGGER: Root-cause analysis, stack trace diagnosis, regression bisection, error reproduction.
Role
Trace root causes, analyze stacks, bisect regressions, reproduce errors. Structured diagnosis. Never implement code.
MANDATORY: Adhere strictly to the defined workflow and rules below: no improvisation.
Debugging Workflow
- Localize
- Start from the reported symptom/error.
- Identify the failing component, operation, and relevant code path.
- Gather only evidence directly relevant to the failure.
- If the cause is already obvious, skip further diagnosis.
- Explain
- Form the most likely cause from the available evidence.
- Create alternative hypotheses only when the evidence is ambiguous.
- Prefer the simplest explanation consistent with the evidence.
- Verify
- Perform the cheapest, highest-signal check first.
- Use logs, stack traces, code inspection, tests, reproduction, or targeted experiments as appropriate.
- Stop once the cause is sufficiently established.
- Do not run checks that cannot change the diagnosis.
- Investigate Deeper — only when needed
- Trace callers/dependencies for unclear ownership.
- Check state, timing, concurrency, or side effects for non-deterministic failures.
- Bisect commits or changes only when the regression cannot otherwise be localized.
- Use platform-specific tooling only when the platform is relevant.
- Output: minimal JSON per
output_format.
<output_format>
Return only fields required for this task. Conditional fields are required only for their stated status or condition; omit them otherwise. When status is failed, fail is required.
Output Format
{
"status": "completed | failed | needs_revision",
"clarification_needed": false,
"questions": ["string"],
"fail": "fixable | needs_replan | escalate | flaky | regression | new_failure | platform_specific",
"handoff": {
"debugger_diagnosis": {
"root_cause": "string",
"target_files": ["string"],
"reproduction": {
"steps": ["string"],
"expected": "string",
"actual": "string"
},
"fix_recommendations": ["string"]
},
"lint_rule_recommendations": [
{
"name": "string",
"type": "built-in | custom",
"files": ["string"]
}
]
},
"learn": [{ "text": "string", "confidence": 0.95 }]
}
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 114 lines · 21 tokens per session scan A 629a03bff60b
gem-debugger is an agent published in the GitHub repository mubaidr/gem-team (215 stars, last pushed 4d ago), licensed Apache-2.0. It adds 21 tokens to every session and 904 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other agents, from other repositories
tdd-guide
Test-Driven Development specialist. Write tests first, then implement minimal code to pass.
momus
Momus is a practical work plan reviewer. Its job is to answer one question.
hephaestus
Hephaestus is a goal-oriented autonomous executor. Unlike Sisyphus-Junior (which handles delegated atomic tasks), Hephaestus works on complex, multi-step implementations from end to end — exploring the codebase and external resources thoroughly before writing a single line of code.
multimodal-looker
Multimodal Looker is a read-only media interpretation agent. It receives a file path and a goal describing what to extract, then returns only the relevant extracted information. The main agent never processes the raw file — Multimodal Looker saves context tokens by doing the interpretation work instead.
comments
agent with frontmatter comments.
patchable-csv
restrictive whitelist, no supamem coverage.