Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add hoatv2211/GameStudio-CodexKIT --skill bug-hunt-swarmgit clone --depth 1 https://github.com/hoatv2211/GameStudio-CodexKITWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/hoatv2211/gamestudio-codexkit/bug-hunt-swarm)<a href="https://agentmods.dev/skills/hoatv2211/gamestudio-codexkit/bug-hunt-swarm"><img src="https://agentmods.dev/badge/skills/hoatv2211/gamestudio-codexkit/bug-hunt-swarm/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/hoatv2211/gamestudio-codexkit/bug-hunt-swarm"><img src="https://agentmods.dev/badge/skills/hoatv2211/gamestudio-codexkit/bug-hunt-swarm.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00042 | $0.00809 |
| Opus 5 | $0.00021 | $0.00404 |
| Sonnet 5 | $0.00008 | $0.00162 |
| Haiku 4.5 | $0.00004 | $0.00081 |
Grade A, and why
bug-hunt-swarm scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 83 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Bug Hunt Swarm
Overview
Explore an unknown failure through independent read-only hypotheses, then converge on the smallest discriminating reproduction or instrumentation plan.
When to use
Use for intermittent crashes, unclear subsystem ownership, multi-service failures, race conditions, protocol mismatches, or symptoms with several plausible root causes.
When NOT to use
Do not use for reviewing an already known patch, implementing fixes, or parallel lanes that would need to mutate the same environment.
Required inputs and context discovery
Collect symptom, expected behavior, frequency, environment, timeline, known-good snapshot, logs, subsystem map, active processes, safe read-only commands, and integrator owner.
Safety and risk level
Lanes are read-only. They may propose instrumentation but cannot patch files, start or stop services, change databases, or modify shared fixtures.
Workflow
- Freeze one symptom statement and shared evidence packet. Completion criterion: all lanes investigate the same observable failure.
- Assign disjoint hypotheses such as data, timing, config, protocol, rendering, or environment. Completion criterion: each lane has a falsifiable question and suspect paths.
- Run independent read-only inspection and reproduction attempts. Completion criterion: each lane returns evidence, counterevidence, and confidence.
- Rank hypotheses by explanatory power and cost of the next experiment. Completion criterion: the integrator selects one discriminating action.
- Hand the selected action to
evidence-first-debuggingfor mutation or instrumentation. Completion criterion: the swarm ends before any lane writes files.
Evidence and output contract
Produce bug packets containing hypothesis, evidence, counterevidence, reproduction status, suspect paths, proposed experiment, confidence, and integrator rank.
Handoff contract
Record symptom, shared snapshot, lanes, commands, ranked hypotheses, unresolved evidence, and the next single experiment with its owner.
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 83 lines · 42 tokens per session scan A e89c66741853
bug-hunt-swarm is a skill published in the GitHub repository hoatv2211/GameStudio-CodexKIT (3 stars, last pushed 2d ago), licensed MIT. It adds 42 tokens to every session and 809 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
vibe-gap-closure-loop
Closes production readiness gaps from a gap analysis document. Autonomously loops dependency analysis, parallel agent dispatch, test gate, re-audit, and repeats until target dimensions reach target score.
vibe-gap-analysis
Assesses production readiness or audits a codebase against its specs. Supports quick static mode, deep 17-dimension audit, or single-dimension focus.
vibe-debugging-journal
Records resolved bugs in a persistent journal — symptom, root cause, difficulty, fix, prevention. Organized by category for future reference.
vibe-production-mindset
Sets quality expectations for implementation — "imagine this serves 1 million users." Checks for observability, error handling, input validation, graceful degradation.
vibe-wave-based-remediation
Systematically fixes large backlogs of issues using prioritized waves. Use when facing 10+ issues, bugs, or gaps to address.
langsmith-tracing
LangSmith tracing and debugging setup for LLM applications. Configure observability, capture traces, and enable debugging for LangChain/LangGraph agents.