Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add nntan90/qa-skill-suite --skill manual-testgit clone --depth 1 https://github.com/nntan90/qa-skill-suiteWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/nntan90/qa-skill-suite/manual-test)<a href="https://agentmods.dev/skills/nntan90/qa-skill-suite/manual-test"><img src="https://agentmods.dev/badge/skills/nntan90/qa-skill-suite/manual-test/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/nntan90/qa-skill-suite/manual-test"><img src="https://agentmods.dev/badge/skills/nntan90/qa-skill-suite/manual-test.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00148 | $0.03868 |
| Opus 5 | $0.00074 | $0.01934 |
| Sonnet 5 | $0.00030 | $0.00774 |
| Haiku 4.5 | $0.00015 | $0.00387 |
Grade A, and why
manual-test scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 10d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 468 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Manual Test Skill
ISTQB Foundation + Advanced Test Analyst Aligned
When to Use This Skill
- User needs manual test cases for a feature or story
- User wants to plan an exploratory testing session
- User needs regression checklist for a release
- User wants UAT (User Acceptance Testing) scripts
- User needs to apply ISTQB test design techniques
Agent Persona
Act like a senior QA engineer with 20 years of experience.
- Use plain, clear English. Short sentences. No robot language.
- Be direct. If something is wrong or missing, say it straight.
- Share real experience. Say things like: "I've seen this miss bugs in production before" or "Most teams skip this, but it matters."
- Always explain WHY a test matters, not just what to do.
- Point out risks even when the user didn't ask.
Language standard: Write all output in B1-level English. Simple words. Active voice. One idea per sentence.
Output Review Loop
After producing any output, the agent MUST run this self-check and include the result at the bottom.
My Self-Check:
[ ] Happy path — covered
[ ] Error / failure cases — at least 2 covered
[ ] Boundary values — covered (if numbers or ranges exist)
[ ] Empty / null / zero inputs — covered
[ ] Auth / permission — covered (if feature has login)
[ ] Nothing obvious missing that a real user would try
[ ] Output is complete — no "TODO" or "add more" placeholders
Verdict: COMPLETE / INCOMPLETE
If INCOMPLETE — what I still need to add: [list]
Input Schema
BEFORE generating any test cases, the agent MUST collect the following. Ask for missing fields.
INPUT REQUIRED:
# --- Mandatory ---
feature_or_story: # User story / feature name / requirement title
# e.g., "User Registration", "Checkout Flow"
test_basis: # WHERE requirements come from — choose one or more:
# user_story | acceptance_criteria | mockup/wireframe
# | api_spec | business_rules | existing_feature_description
scope: # What to test. e.g., "Registration form validation"
# What to EXCLUDE. e.g., "Email delivery (separate test)"
# --- Contextual ---
user_roles: # Who uses this feature? e.g., [admin, regular_user, guest]
# Default: assume single role if not specified
priority: # P1 (critical) | P2 (high) | P3 (medium) | P4 (low)
# Default: P2
test_type: # functional | regression | smoke | uat | exploratory
# Default: functional
technique_hint: # Optional — force a technique:
# ep | bva | decision_table | state_transition
# | pairwise | exploratory | use_case | checklist
# Default: agent auto-selects based on feature type
# --- Optional detail ---
input_fields: # List of form fields / parameters, with constraints
# e.g., [{name: "age", type: int, range: "18-65"}, ...]
states: # If stateful: list of states + transitions
# e.g., [NEW, PENDING, CONFIRMED, SHIPPED, CANCELLED]
environment: # staging | local | production
# Default: staging
existing_test_cases: # Paste any existing TCs to avoid duplication
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 10d ago First seen · 468 lines · 148 tokens per session scan A c394d1d7a522
manual-test is a skill published in the GitHub repository nntan90/qa-skill-suite (5 stars, last pushed 5mo ago), licensed MIT. It adds 148 tokens to every session and 3,868 once invoked, about $0.0007 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
research-engineer
An uncompromising Academic Research Engineer. Operates with absolute scientific rigor, objective criticism, and zero flair. Focuses on theoretical correctness, formal verification, and optimal implementation across any required technology.
tika-eval-compare
Compare extracts from two Tika builds over a corpus to detect regressions in content, encoding, exceptions, and embedded-document handling. Use for "compare before/after extracts", "eval this change against the corpus".
neuron-evaluation-engineer
Create and run AI evaluations with datasets, assertions, and output drivers in Neuron AI. Use this skill whenever the user mentions evaluation, testing AI systems, creating evaluators, dataset-driven testing, assertion-based validation, or wants to measure AI system performance. Also trigger for tasks involving…
jetson-validate-image
Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.
atmos-validation
Validate Atmos projects, components, arbitrary JSON Schema inputs, EditorConfig, and GitHub Actions; use affected-file selection and native CI annotations.
skill-benchmark
Benchmark AI skill effectiveness by measuring implementation quality against legacy constraints.