Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add regiellis/godot-mcp-go --skill swallowtail-qagit clone --depth 1 https://github.com/regiellis/godot-mcp-goWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/regiellis/godot-mcp-go/swallowtail-qa)<a href="https://agentmods.dev/skills/regiellis/godot-mcp-go/swallowtail-qa"><img src="https://agentmods.dev/badge/skills/regiellis/godot-mcp-go/swallowtail-qa/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/regiellis/godot-mcp-go/swallowtail-qa"><img src="https://agentmods.dev/badge/skills/regiellis/godot-mcp-go/swallowtail-qa.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00050 | $0.00959 |
| Opus 5.5 | $0.00020 | $0.00384 |
| Sonnet 5.5 | $0.00010 | $0.00192 |
| Haiku 4.5 | $0.00005 | $0.00096 |
Grade A, and why
swallowtail-qa scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 14d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 31 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Swallowtail Game QA
Use the local swallowtail qa commands. Read the command and extension contract before creating a scenario. The CLI embeds the runner, mascot, fonts and PDF template. Python 3.10+ is required; PDF generation additionally needs ReportLab 4 (SWALLOWTAIL_PYTHON selects the interpreter). Do not install dependencies globally without considering the project's environment.
Establish the release target, build provenance, intended platforms, representative gameplay, and explicit acceptance criteria from the request and repository. Progress with available evidence; ask only for material missing information. Export locally when authorized. Publishing or store uploads are separate actions.
Run against isolated test saves. The worker isolates Windows APPDATA/LOCALAPPDATA and Linux XDG directories. Check games using custom paths, macOS saves or cloud synchronization before assuming isolation. Retain original saves and fixtures.
Keep source and packaged evidence distinct. qa run source mode injects a SceneTree runner and calls the game's GDScript run(qa) scenario. Package mode launches the exact executable without script injection. Hardened exports can ignore --script or refuse --path; an exit-zero launch is not proof an injected test ran. Missing scenario receipts fail source runs. Never put QA addons into a player build solely to get a green test.
Choose checks that cover the game's actual release risks: normal boot, repeated transitions, representative gameplay and restart, input/focus, pause/settings, save/reload, clean exit. Extend a small game-owned GDScript scenario with qa.check, qa.phase, qa.wait, and qa.screenshot; reuse the project's existing fixtures and custom-command logic where appropriate. Editor-only mcp_commands/*.gd remain editor extensions, not packaged tests.
For timing, separate cold boot, first entry, repeated entries, steady gameplay, and teardown. Keep capture resolution, renderer, vsync, power state and hardware recorded. PresentMon is optional Windows capture; process-frame timing is a different measurement. Name screenshot phases separately because reading back pixels can stall a frame. Keep raw CSVs, including spikes. Do not infer CPU/GPU causes from frame duration alone; use source profiler attribution as a separate experiment.
Treat the full process log as evidence, including diagnostics printed after a PASS receipt and during shutdown. A timeout, engine error, missing receipt or missing requested capture is a failure. Pending manual checks mean incomplete coverage. Use qa attest only with an actual observed or user-reported result, named check and substantive evidence. Never fabricate visual, audio, controller, accessibility or hardware coverage.
Before an operator-driven capture, establish that the operator can see and interact with the exact game window. A Ready response before launch, a live PID, a boot log or presentation events do not prove visibility. If the window is unavailable, retain process data with that limitation and leave manual checks pending; do not label the trace gameplay. A later operator correction supersedes an earlier form response and must be recorded in the run and regenerated report.
Compare only compatible runs with qa compare; environment or scenario changes make a baseline incompatible. Explain differences rather than relaxing matching to force a result. Use explicit game-specific budgets; a single machine does not establish minimum hardware support. Retain the approved baseline as an immutable run directory and compare new runs to it.
Generate the branded PDF with qa report, inspect rendered pages, and verify that screenshots, labels, hashes, findings and limits match the raw evidence. Append follow-up results to an existing launch report when requested; preserve historical failures as resolved history rather than deleting them. Summarize actionable findings with reproduction, severity, evidence, owning repo/file, fix and retest status.
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 14d ago First seen · 31 lines · 50 tokens per session scan A 32603c29c563
swallowtail-qa is a skill published in the GitHub repository regiellis/godot-mcp-go (68 stars, last pushed 2d ago), licensed MIT. It adds 50 tokens to every session and 959 once invoked, about $0.0002 per session on Opus 5.5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-27.
Other skills, from other repositories
smoke-check
Run the critical path smoke test gate before QA hand-off. Executes the automated test suite, verifies core functionality, and produces a PASS/FAIL report. Run after a sprint's stories are implemented and before manual QA begins. A failed smoke check means the build is not ready for QA.
godot-testing-qa
Domain skill — Run a sequence of runtime actions and assertions. Assert runtime node existence and properties. Assert visible runtime text. Compare two PNGs using bounded pixel sampling. Sample runtime performance for a bounded frame count. Return the latest runtime test report.
unity-ui
A Unity user-interface development and testing workflow for UI Toolkit and uGUI, Unity's two main UI systems.
e2e-testing
Guide for running end-to-end tests of the Qwen Code CLI, including headless mode, MCP server testing, and API traffic inspection. Use this skill whenever you need to verify CLI behavior with real model calls, reproduce user-reported bugs end-to-end, test MCP tool integrations, or inspect raw API request/response…
terminal-capture
Automates terminal UI screenshot testing for CLI commands. Applies when reviewing PRs that affect CLI output, testing slash commands (/about, /context, /auth, /export), generating visual documentation, or when 'terminal screenshot', 'CLI test', 'visual test', or 'terminal-capture' is mentioned.
agent-reproduce-align
Use after a Codex or Claude Code feature has been implemented in Qwen Code to run the selected reference agent and Qwen Code under the same scenario, capture HTTP and terminal traces, compare request bodies, tool/function schemas, outputs, and iterate until the reproduced behavior is close enough.