reviewing-game-qa

reviewing-game-qa is a skill for Claude Code from Bulugulu/game-qa. It costs 74 tokens per session (942 once invoked), scanned A, original, MIT.

A browser-game quality check that reviews recorded test evidence and decides whether a feature works for players.

In plain words
What is it for?
Reviewing assigned game test cases, checking screenshots and other captured evidence, deciding pass or fail, and filing tickets as findings appear.
Why use it?
It catches problems that automated checks may miss, such as a game feeling broken even when its internal values look correct. Missing evidence is reported as unverified instead of guessed.

Skill for Claude Code

Written for Claude Code: shipped in a Claude Code plugin. Also seen: mentions subagents.

Part of the game-qa plugin — 6 skills shipped together

Good fit Reviewing assigned game test cases, checking screenshots and other captured evidence, deciding pass or fail, and filing tickets as findings appear.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/bulugulu/game-qa/reviewing-game-qa
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add Bulugulu/game-qa --skill reviewing-game-qa
Clone the repo
git clone --depth 1 https://github.com/Bulugulu/game-qa

Made for: Claude Code.

Or install game-qa, the plugin that ships this one along with the rest of its 6 skills.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for reviewing-game-qa

README.md
[![agentmods](https://agentmods.dev/badge/skills/bulugulu/game-qa/reviewing-game-qa/github.svg)](https://agentmods.dev/skills/bulugulu/game-qa/reviewing-game-qa)
Your own site
<a href="https://agentmods.dev/skills/bulugulu/game-qa/reviewing-game-qa"><img src="https://agentmods.dev/badge/skills/bulugulu/game-qa/reviewing-game-qa/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for reviewing-game-qa

Your own site · 80×15
<a href="https://agentmods.dev/skills/bulugulu/game-qa/reviewing-game-qa"><img src="https://agentmods.dev/badge/skills/bulugulu/game-qa/reviewing-game-qa.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 74 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 942 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00074 $0.00942
Opus 5 $0.00037 $0.00471
Sonnet 5 $0.00015 $0.00188
Haiku 4.5 $0.00007 $0.00094

Measured 11d ago against content hash fca393f44d8a, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-11, from the pricing page.

Security

Grade A, and why

reviewing-game-qa scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/game-qa/skills/reviewing-game-qa/SKILL.md · 83 lines

How it starts

The opening of the file, as written. The whole thing — 83 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Reviewing Game QA

You are the QA reviewer. The QA Lead assigned you cases. Engineers ran journeys and produced evidence bundles. Your job: own the cases, judge each from the evidence, file tickets as you go.

You're the only one who writes verdicts. The script proposes; you dispose.

Core principle: Mechanical pass + feels broken to a player = fail. Always.

The Iron Law

GAME FEATURES ONLY.
CASES ARE YOURS. EVIDENCE IS YOUR INPUT. TICKETS ARE YOUR OUTPUT.
NEVER COMPROMISE: INSUFFICIENT EVIDENCE = UNVERIFIED, NOT GUESS.

The Loop

For each case in your slice:

  1. Read the case spec. playerFlow, evidenceRequirements, playerReviewQuestions, inputMode, reviewMode.
  2. Locate the evidence in journey output(s). Verify the journey actually captured what the case requires.
  3. Validate mechanical checks against state samples (e.g., hp===0, score incremented by N, flag toggled, etc.).
  4. Answer every playerReviewQuestion from screenshots / HUD crops / DOM snippets / console / timings. Don't skip — these are the player-perspective gates the script can't see.
  5. Mark verdict:
    • pass — mechanical and visual both match
    • softPass — mechanical passes, minor visual flake or duplicate-of-unit-coverage
    • fail — anything that wouldn't ship
    • unverified-pending-coverage — evidence missing; need additional capture
    • unverified-pending-tooling — capability-gap blocked the case
    • unverified-blocked — upstream journey crash
  6. File tickets immediately on any finding. Don't batch — the engineer is fixing in parallel.
  7. Write <reviewer>-<case-id>.json to <artifactDir>/reviews/.

What to Produce

  • One review JSON per case (verdict + cited evidence + observations)
  • Tickets streamed live to <artifactDir>/tickets/ (per references/ticket-schema.md)

Return to Lead: one paragraph — case count, pass/soft/fail/unverified breakdown, tickets filed.

Hard Rules

  • GAME features only. Refuse and route otherwise.
  • Mechanical pass + feels broken = fail. State-correct journey is not a passing case if visuals/HUD/affordance failed.
  • Answer every playerReviewQuestion. Skip = unverified-pending-coverage.
  • File tickets immediately. Engineer needs them streaming, not batched.
  • Insufficient evidence → request more, don't guess. File a capture request via ticket; mark case unverified.
  • No sub-subagents. You are a leaf.

Read the full file on GitHub · 83 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 11d ago First seen · 83 lines · 74 tokens per session scan A fca393f44d8a

Subscribe to this mod's changes

reviewing-game-qa is a skill published in the GitHub repository Bulugulu/game-qa (1 stars, last pushed 4mo ago), licensed MIT. It adds 74 tokens to every session and 942 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

smoke-check

Run the critical path smoke test gate before QA hand-off. Executes the automated test suite, verifies core functionality, and produces a PASS/FAIL report. Run after a sprint's stories are implemented and before manual QA begins. A failed smoke check means the build is not ready for QA.

striderZA/OpenCodeGameStudios · 61 tokens

hearth-playtest

Let the engine hunt bugs for you — bot playtesting via hearth sweep. Seeded bot policies (mash/idle/wander/seek) play a scene headlessly across many seeds and report softlocks, crashes, stuck states, and unmet objectives as a compact evidence report; objectives double as executable acceptance criteria; a failing seed…

echoo19/hearth · 112 tokens

game-playtesting

Test a game end to end like a player across boot, controls, gameplay loops, UI, save state, failure recovery, rendering, audio, performance, and platform behavior. Use after game implementation or for regression investigation.

metaspartan/cybara · 48 tokens

gm-evaluate

Evaluate the current tag's quality: enforce the playable-closed-loop gate, maintain a single cross-tag e2e/ suite that always reflects the current game (add tests for new mechanics, prune tests for mechanics this tag deliberately removed), and reason about gameplay quality. Independent from the build process — fresh…

RandallLiuXin/GodotMaker · 84 tokens

godot-e2e

Write and run E2E (end-to-end) game tests using the godot-e2e framework. Python controls a live Godot game over TCP — Locator-based semantic queries, expect() auto-retry assertions, and engine log capture make failures self-diagnosing. Use this skill whenever you need to: Test actual gameplay: player movement…

RandallLiuXin/GodotMaker · 189 tokens

screenshot

Capture gameplay screenshots using godot-e2e for visual verification. Use when you need to: take a screenshot of the running game, capture multiple screenshots during an E2E scenario, generate reference.png, visually verify game state, or provide screenshots for VQA analysis. Triggers: "screenshot", "capture…

RandallLiuXin/GodotMaker · 87 tokens