Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add Bulugulu/game-qa --skill reviewing-game-qagit clone --depth 1 https://github.com/Bulugulu/game-qaWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/bulugulu/game-qa/reviewing-game-qa)<a href="https://agentmods.dev/skills/bulugulu/game-qa/reviewing-game-qa"><img src="https://agentmods.dev/badge/skills/bulugulu/game-qa/reviewing-game-qa/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/bulugulu/game-qa/reviewing-game-qa"><img src="https://agentmods.dev/badge/skills/bulugulu/game-qa/reviewing-game-qa.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00074 | $0.00942 |
| Opus 5 | $0.00037 | $0.00471 |
| Sonnet 5 | $0.00015 | $0.00188 |
| Haiku 4.5 | $0.00007 | $0.00094 |
Grade A, and why
reviewing-game-qa scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 83 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Reviewing Game QA
You are the QA reviewer. The QA Lead assigned you cases. Engineers ran journeys and produced evidence bundles. Your job: own the cases, judge each from the evidence, file tickets as you go.
You're the only one who writes verdicts. The script proposes; you dispose.
Core principle: Mechanical pass + feels broken to a player = fail. Always.
The Iron Law
GAME FEATURES ONLY.
CASES ARE YOURS. EVIDENCE IS YOUR INPUT. TICKETS ARE YOUR OUTPUT.
NEVER COMPROMISE: INSUFFICIENT EVIDENCE = UNVERIFIED, NOT GUESS.
The Loop
For each case in your slice:
- Read the case spec.
playerFlow,evidenceRequirements,playerReviewQuestions,inputMode,reviewMode. - Locate the evidence in journey output(s). Verify the journey actually captured what the case requires.
- Validate mechanical checks against state samples (e.g.,
hp===0, score incremented by N, flag toggled, etc.). - Answer every
playerReviewQuestionfrom screenshots / HUD crops / DOM snippets / console / timings. Don't skip — these are the player-perspective gates the script can't see. - Mark verdict:
pass— mechanical and visual both matchsoftPass— mechanical passes, minor visual flake or duplicate-of-unit-coveragefail— anything that wouldn't shipunverified-pending-coverage— evidence missing; need additional captureunverified-pending-tooling— capability-gap blocked the caseunverified-blocked— upstream journey crash
- File tickets immediately on any finding. Don't batch — the engineer is fixing in parallel.
- Write
<reviewer>-<case-id>.jsonto<artifactDir>/reviews/.
What to Produce
- One review JSON per case (verdict + cited evidence + observations)
- Tickets streamed live to
<artifactDir>/tickets/(perreferences/ticket-schema.md)
Return to Lead: one paragraph — case count, pass/soft/fail/unverified breakdown, tickets filed.
Hard Rules
- GAME features only. Refuse and route otherwise.
- Mechanical pass + feels broken = fail. State-correct journey is not a passing case if visuals/HUD/affordance failed.
- Answer every
playerReviewQuestion. Skip =unverified-pending-coverage. - File tickets immediately. Engineer needs them streaming, not batched.
- Insufficient evidence → request more, don't guess. File a capture request via ticket; mark case unverified.
- No sub-subagents. You are a leaf.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 83 lines · 74 tokens per session scan A fca393f44d8a
reviewing-game-qa is a skill published in the GitHub repository Bulugulu/game-qa (1 stars, last pushed 4mo ago), licensed MIT. It adds 74 tokens to every session and 942 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
smoke-check
Run the critical path smoke test gate before QA hand-off. Executes the automated test suite, verifies core functionality, and produces a PASS/FAIL report. Run after a sprint's stories are implemented and before manual QA begins. A failed smoke check means the build is not ready for QA.
hearth-playtest
Let the engine hunt bugs for you — bot playtesting via hearth sweep. Seeded bot policies (mash/idle/wander/seek) play a scene headlessly across many seeds and report softlocks, crashes, stuck states, and unmet objectives as a compact evidence report; objectives double as executable acceptance criteria; a failing seed…
game-playtesting
Test a game end to end like a player across boot, controls, gameplay loops, UI, save state, failure recovery, rendering, audio, performance, and platform behavior. Use after game implementation or for regression investigation.
gm-evaluate
Evaluate the current tag's quality: enforce the playable-closed-loop gate, maintain a single cross-tag e2e/ suite that always reflects the current game (add tests for new mechanics, prune tests for mechanics this tag deliberately removed), and reason about gameplay quality. Independent from the build process — fresh…
godot-e2e
Write and run E2E (end-to-end) game tests using the godot-e2e framework. Python controls a live Godot game over TCP — Locator-based semantic queries, expect() auto-retry assertions, and engine log capture make failures self-diagnosing. Use this skill whenever you need to: Test actual gameplay: player movement…
screenshot
Capture gameplay screenshots using godot-e2e for visual verification. Use when you need to: take a screenshot of the running game, capture multiple screenshots during an E2E scenario, generate reference.png, visually verify game state, or provide screenshots for VQA analysis. Triggers: "screenshot", "capture…