windows-qa-engineer

A guide for testing Windows 11 desktop applications by controlling their windows and interface elements. It uses UI automation to interact with the app like a person would.

In plain words
What is it for?
Use it for smoke tests, checking user-interface controls, filling forms, clicking buttons, and verifying dialogs.
Why use it?
It provides a repeatable way to check buttons, forms, dialogs, and other visible parts of a Windows app.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/codealive-ai/ai-driven-development/windows-qa-engineer
Any agent
npx skills add CodeAlive-AI/ai-driven-development --skill windows-qa-engineer
Clone the repo
git clone --depth 1 https://github.com/CodeAlive-AI/ai-driven-development

Made for: Claude Code, Codex.

Per session 107 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,315 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00107 $0.01315
Opus 5 $0.00053 $0.00658
Sonnet 5 $0.00021 $0.00263
Haiku 4.5 $0.00011 $0.00131

Measured 2d ago against content hash a493112192df, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

windows-qa-engineer scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

The scan reads SKILL.md. This mod also ships 3 executable files (scripts/doctor.ps1, scripts/skill_installer.py, scripts/ufo_windows_qa_mcp_server.py), listed below but not scanned — reading those needs a real analyzer, not pattern matching.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/windows-qa-engineer/SKILL.md · 105 lines

How it starts

The opening of the file, as written. The whole thing — 105 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Windows QA Engineer (UFO-powered)

You are an AI-QA operator on the SAME Windows 11 desktop as the SUT. All automation uses UFO's real MCP tools (UICollector, HostUIExecutor, AppUIExecutor) -- no mocks.

Auto-Setup (when MCP tools are missing)

If UFO tools are NOT available as MCP tools, run setup before QA work:

  1. Run: python "<skill-dir>/scripts/skill_installer.py" --project-dir "<project-root>"
  2. Parse the JSON output — if success is true, tell user to restart Claude Code
  3. If failed, show the error and direct user to references/setup.md for manual install

Mandatory Workflow

Follow this sequence for every test run. Do not skip steps.

1. Discover windows

  • Call qa_refresh_and_list_windows()
  • Identify the SUT window by title hint from the user

2. Select window

  • Call select_application_window(id, name) (HostUIExecutor)
  • Call capture_window_screenshot() (UICollector) -- baseline screenshot

3. Collect controls

  • Call qa_refresh_controls(field_list=["label","control_text","control_type","automation_id","control_rect"])
  • Anchor on id + control_text / automation_id when the returned tree is usable
  • If control collection returns an error or an empty tree for a large/legacy WinForms window, continue with screenshot inspection and coordinate actions; do not repeatedly force full UIA subtree scans

4. Interact

  • Use click_input(id, name), set_edit_text(id, name, text), keyboard_input(id, name, keys)
  • Coordinate actions only as last resort (document why)
  • Re-collect controls after navigation or dialog open

5. Assert

  • Read with texts(id, name) and compare against expected
  • Prefer qa_wait_for_text_contains(id, name, expected, timeout_s=10) over sleeps
  • Screenshot after each major checkpoint

6. Report

  • Fill assets/test-case.md template
  • Numbered execution log (step -> tool call -> result)
  • Final PASS/FAIL with exact failing assertion if applicable
  • Attach screenshot base64 strings from capture_window_screenshot()

Read the full file on GitHub · 105 lines

Files

What ships with it

8 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 105 lines · 107 tokens per session scan A a493112192df

Subscribe to this mod's changes

windows-qa-engineer is a skill published in the GitHub repository CodeAlive-AI/ai-driven-development (131 stars, last pushed 4d ago), licensed MIT. It adds 107 tokens to every session and 1,315 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

map-debug

Structured MAP debugging via task-decomposer, actor, and monitor agents. Use when reproducing a bug, isolating a regression, or diagnosing an error with specialized agents — including failing or flaky tests (pytest AssertionError), crashes and segmentation faults, memory-corruption or memory errors in native/C…

azalio/map-framework · 213 tokens

map-learn

Capture reusable lessons after a completed MAP workflow. Use when a MAP run has finished and you want rules written to .claude/rules/learned/ from a workflow summary or handoff. Do NOT use during active implementation.

azalio/map-framework · 51 tokens

map-task

Execute a single subtask from an existing MAP plan via Actor and Monitor. Use when map-plan has decomposed work and you want fine-grained control over one subtask. Do NOT use without an existing plan; run map-plan first.

azalio/map-framework · 51 tokens

map-explain

Deep walkthrough of code, a diff, or the whole project — problem, entities, flow, load-bearing-line rationale, side effects, assumptions, breakage. Use when learning unfamiliar code or auditing a diff.

azalio/map-framework · 46 tokens

map-auto

Single-entry autonomous autopilot: routes a task through the existing MAP workflows via routetask, then drives the selected chain (map-plan -> map-efficient -> map-check -> map-review, as routed) end-to-end to a committed feature branch in one session, auto-approving routine workflow-control holds and hard-stopping on…

azalio/map-framework · 165 tokens

map-resume

Resume an interrupted MAP workflow from .map/ /stepstate.json checkpoint. Use when returning after context exhaustion, /clear, or a session crash mid-workflow. Do NOT use to start new work; use map-plan or map-efficient.

azalio/map-framework · 53 tokens