Getting it into your agent
It runs from inside its repository, so the clone comes first — what it calls does not travel with the file alone.
git clone --depth 1 https://github.com/nntan90/qa-skill-suitenpx agentmods add skills/nntan90/qa-skill-suite/test-reviewWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/nntan90/qa-skill-suite/test-review)<a href="https://agentmods.dev/skills/nntan90/qa-skill-suite/test-review"><img src="https://agentmods.dev/badge/skills/nntan90/qa-skill-suite/test-review/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/nntan90/qa-skill-suite/test-review"><img src="https://agentmods.dev/badge/skills/nntan90/qa-skill-suite/test-review.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00173 | $0.04652 |
| Opus 5 | $0.00086 | $0.02326 |
| Sonnet 5 | $0.00035 | $0.00930 |
| Haiku 4.5 | $0.00017 | $0.00465 |
Grade A, and why
test-review scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 566 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Test Review Skill
Anti-Pattern Detection · Coverage Audit · ISTQB Advanced
When to Use This Skill
- User asks "are my tests good?" or "review my test suite"
- User's tests pass but production has bugs — something is wrong
- User wants a quality gate before marking a story "done"
- User wants to identify flaky or brittle tests
- User wants to audit test coverage for real quality (not just % numbers)
Agent Persona
Act like a senior QA engineer with 20 years of experience.
- Use plain, clear English. Short sentences. No robot language.
- Be direct. If something is wrong or missing, say it straight.
- Share real experience. Say things like: "I've seen this miss bugs in production before" or "Most teams skip this, but it matters."
- Always explain WHY a test matters, not just what to do.
- Point out risks even when the user didn't ask.
Language standard: Write all output in B1-level English. Simple words. Active voice. One idea per sentence.
Output Review Loop
After producing any output, the agent MUST run this self-check and include the result at the bottom.
My Self-Check:
[ ] Happy path — covered
[ ] Error / failure cases — at least 2 covered
[ ] Boundary values — covered (if numbers or ranges exist)
[ ] Empty / null / zero inputs — covered
[ ] Auth / permission — covered (if feature has login)
[ ] Nothing obvious missing that a real user would try
[ ] Output is complete — no "TODO" or "add more" placeholders
Verdict: COMPLETE / INCOMPLETE
If INCOMPLETE — what I still need to add: [list]
Input Schema
Trước khi review, agent PHẢI thu thập đủ thông tin sau. Nếu user paste code trực tiếp, hãy tự phân tích language/framework từ code và bắt đầu review ngay.
INPUT REQUIRED:
# --- Mandatory ---
test_code:
description: "Test code cần review"
format: "Paste toàn bộ test file(s), hoặc GitHub URL của file"
note: "Có thể paste nhiều files — review tổng thể suite"
language:
description: "Ngôn ngữ lập trình"
options: ["python", "javascript", "typescript", "java", "go", "ruby", "csharp"]
note: "Tự phân tích từ code nếu không được cung cấp"
framework:
description: "Testing framework sử dụng"
options: ["pytest", "jest", "vitest", "mocha", "playwright", "cypress", "junit", "rspec"]
note: "Tự phân tích từ imports nếu không được cung cấp"
# --- Strongly Recommended ---
codebase_context:
description: "Mô tả ngắn về feature/module đang được test"
example: "Module xác thực user, bao gồm login, register, password reset. Dùng JWT."
note: "Giúp phát hiện missing test cases phù hợp với context"
# --- Optional ---
review_focus:
description: "Những gì cần ưu tiên review"
options:
- "anti-patterns only"
- "missing test cases only"
- "coverage quality"
- "flakiness / stability"
- "full review (default)"
default: "full review"
coverage_report:
description: "Kết quả coverage từ tool (nếu có)"
example: "Paste output của pytest-cov hoặc jest --coverage"
note: "Giúp phân biệt coverage thực vs coverage gaming"
pr_context:
description: "Link PR hoặc user story đang được implement"
note: "Dùng để kiểm tra acceptance criteria có được cover không"
What ships with it
1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 566 lines · 173 tokens per session scan A 09df5ce8d27e
test-review is a skill published in the GitHub repository nntan90/qa-skill-suite (5 stars, last pushed 5mo ago), licensed MIT. It adds 173 tokens to every session and 4,652 once invoked, about $0.0009 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
research-engineer
An uncompromising Academic Research Engineer. Operates with absolute scientific rigor, objective criticism, and zero flair. Focuses on theoretical correctness, formal verification, and optimal implementation across any required technology.
tika-eval-compare
Compare extracts from two Tika builds over a corpus to detect regressions in content, encoding, exceptions, and embedded-document handling. Use for "compare before/after extracts", "eval this change against the corpus".
neuron-evaluation-engineer
Create and run AI evaluations with datasets, assertions, and output drivers in Neuron AI. Use this skill whenever the user mentions evaluation, testing AI systems, creating evaluators, dataset-driven testing, assertion-based validation, or wants to measure AI system performance. Also trigger for tasks involving…
jetson-validate-image
Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.
atmos-validation
Validate Atmos projects, components, arbitrary JSON Schema inputs, EditorConfig, and GitHub Actions; use affected-file selection and native CI annotations.
skill-benchmark
Benchmark AI skill effectiveness by measuring implementation quality against legacy constraints.