windows-qa-engineer

$npx mdskill add CodeAlive-AI/ai-driven-development/windows-qa-engineer

Automatically discovers and selects the Windows app for testing

  • Identifies the target application window based on user-provided title hint.
  • Deploys UFO's HostUIExecutor to select the correct application window.
  • Uses UICollector to capture a baseline screenshot of the identified window.
  • Provides the necessary information and control for subsequent automation steps.

SKILL.md

.github/skills/windows-qa-engineerView on GitHub ↗
---
name: windows-qa-engineer
description: Use when testing Windows 11 desktop apps (WinForms/WPF/UWP) via UFO UIA/Win32 automation MCP. Triggers on "test this Windows app", "QA the app", "run smoke test", "click the button", "fill the form", "check the UI", "Windows automation", "UFO QA", "verify the dialog", or any Windows desktop UI testing task. Not for web/browser testing (use Playwright), mobile testing, or non-Windows platforms.
metadata:
  compatibility: Windows 11, Python 3.10+, UFO (github.com/microsoft/UFO), fastmcp
---

# Windows QA Engineer (UFO-powered)

You are an AI-QA operator on the SAME Windows 11 desktop as the SUT.
All automation uses UFO's real MCP tools (UICollector, HostUIExecutor, AppUIExecutor) -- no mocks.

## Auto-Setup (when MCP tools are missing)

If UFO tools are NOT available as MCP tools, run setup before QA work:

1. Run: `python "<skill-dir>/scripts/skill_installer.py" --project-dir "<project-root>"`
2. Parse the JSON output — if `success` is true, tell user to restart Claude Code
3. If failed, show the error and direct user to [references/setup.md](references/setup.md) for manual install

## Mandatory Workflow

Follow this sequence for every test run. Do not skip steps.

### 1. Discover windows
- Call `qa_refresh_and_list_windows()`
- Identify the SUT window by title hint from the user

### 2. Select window
- Call `select_application_window(id, name)` (HostUIExecutor)
- Call `capture_window_screenshot()` (UICollector) -- baseline screenshot

### 3. Collect controls
- Call `qa_refresh_controls(field_list=["label","control_text","control_type","automation_id","control_rect"])`
- Anchor on `id` + `control_text` / `automation_id` when the returned tree is usable
- If control collection returns an error or an empty tree for a large/legacy WinForms window, continue with screenshot inspection and coordinate actions; do not repeatedly force full UIA subtree scans

### 4. Interact
- Use `click_input(id, name)`, `set_edit_text(id, name, text)`, `keyboard_input(id, name, keys)`
- Coordinate actions only as last resort (document why)
- Re-collect controls after navigation or dialog open

### 5. Assert
- Read with `texts(id, name)` and compare against expected
- Prefer `qa_wait_for_text_contains(id, name, expected, timeout_s=10)` over sleeps
- Screenshot after each major checkpoint

### 6. Report
- Fill [assets/test-case.md](assets/test-case.md) template
- Numbered execution log (step -> tool call -> result)
- Final PASS/FAIL with exact failing assertion if applicable
- Attach screenshot base64 strings from `capture_window_screenshot()`

## Tool Reference

| Tool | Server | Purpose |
|------|--------|---------|
| `qa_refresh_and_list_windows` | QA helper | Refresh + list all windows |
| `select_application_window` | HostUIExecutor | Select SUT by id+name |
| `get_app_window_controls_info` | UICollector | Raw control tree; use only when helper output is insufficient |
| `capture_window_screenshot` | UICollector | Screenshot selected window |
| `click_input` | AppUIExecutor | Click control by id+name |
| `set_edit_text` | AppUIExecutor | Type into control |
| `keyboard_input` | AppUIExecutor | Send keystrokes |
| `texts` | AppUIExecutor | Read control text |
| `qa_wait_for_text_contains` | QA helper | Poll until text matches |
| `qa_refresh_controls` | QA helper | Re-collect control tree with fail-soft parsing |

## Example: Login Smoke Test

User says: "Test the login flow on MyApp"

```
1. qa_refresh_and_list_windows() → find "MyApp - Login"
2. select_application_window(id="3", name="MyApp - Login")
3. capture_window_screenshot() → baseline
4. qa_refresh_controls(field_list=["label","control_text","control_type","automation_id","control_rect"])
   → find username (id=12), password (id=14), login button (id=16)
5. set_edit_text(id="12", name="Username", text="testuser")
6. set_edit_text(id="14", name="Password", text="pass123")
7. click_input(id="16", name="Login")
8. qa_wait_for_text_contains(id="20", name="WelcomeLabel", expected_substring="Welcome", timeout_s=10)
   → {"ok": true, "text": "Welcome, testuser"}
9. capture_window_screenshot() → post-login
10. Report: PASS
```

## Error Handling

**No windows found**: Re-check the SUT is running. Call `qa_refresh_and_list_windows()` again. If still empty, ask the user to confirm the app is open.

**Empty control tree**: The window may not have finished loading. Wait 2-3 seconds, then `qa_refresh_controls(field_list=[...])`. If still empty, try `CONTROL_BACKEND=win32` (see setup.md). For large or legacy WinForms apps, avoid repeated full UIA subtree scans and use screenshot plus targeted coordinates.

**Control not clickable / action fails**: Re-collect controls (the tree may have changed after navigation). If the control lacks a usable id, fall back to coordinate-based action and document why.

**MCP tools not found**: Run auto-setup first (see [Auto-Setup](#auto-setup-when-mcp-tools-are-missing) above). If auto-setup fails, direct the user to [references/setup.md](references/setup.md) and run `doctor.ps1`.

## Detailed Workflows

See [references/qa-workflows.md](references/qa-workflows.md) for more examples, locator strategy, and common patterns.

## Setup

See [references/setup.md](references/setup.md) for UFO installation, MCP configuration, and diagnostics.

More from CodeAlive-AI/ai-driven-development

SkillDescription
agentic-readinessAudit and improve repositories for reliable agentic work across Codex and Codex App, Claude Code, and OpenCode. Use when reviewing AGENTS.md or CLAUDE.md quality and discovery, instruction routing in monorepos or meta-repos, agent settings, MCP configuration, skills, subagents, context budgets, or repository organization for coding agents.
agents-consiliumQuery external AI agents (Codex, Gemini, OpenCode, Claude Code headless) in parallel for independent second opinions, code review, bug investigation, and consensus on high-stakes decisions. Agents and models are configurable in config.json. Use for architecture choices, security review, or ambiguous problems where independent perspectives matter. Not for simple questions answerable from docs or the codebase — use web search or repo exploration instead.
bug-fix-protocol8-step disciplined bug-fix protocol that treats every production bug as two failures — the code defect itself and the testing system that allowed it through. Use when fixing a production bug, investigating a regression, writing a post-mortem, or auditing a missed defect. Triggers on "fix this bug", "production bug", "regression test", "post-mortem", "test gap", "why did the tests miss this".
code-that-fits-in-your-headSoftware engineering heuristics from Mark Seemann's book Code That Fits in Your Head (2021). Use when writing new code, reviewing code, refactoring, designing APIs, handling validation and invariants, writing unit tests, debugging defects, performing security review (STRIDE), or setting up a new code base. Covers decomposition (cyclomatic complexity, 80/24 rule, cohesion, fractal architecture), encapsulation (invariants, parse-don't-validate, Postel's law), outside-in TDD (walking skeleton, AAA, triangulation, devil's advocate), API design (affordance, poka-yoke, CQS), git/PR hygiene (50/72 commits, small commits, code review), feature flags, Strangler pattern, bisection debugging, logging with decorators, and STRIDE threat modelling. Not for language-specific syntax, framework tutorials, production incident response, or performance profiling.
fetch-url-as-markdownFetch a web page (URL) and return clean Markdown via local trafilatura, with Exa MCP as a fallback for JS-rendered or anti-bot pages. Use when the user asks to read, fetch, scrape, summarize, or quote a URL — prefer this over the built-in WebFetch tool. Don't use for binary files (PDFs, images, archives) or for fetching API/JSON endpoints.
fpf-problem-solvingFirst Principles Framework (FPF) — thinking amplifier. Use when user wants to think through a complex problem, architect a system, evaluate alternatives, decompose complexity, classify problems, define quality attributes, plan rigorously, apply an FPF pattern to a first useful result, decide under uncertainty, establish causality, reason about time and trends, describe or synthesize architecture, check mathematical model fit, distinguish relation kinds or occurrences, govern ontic/U-kind admission, publish multi-view artifacts, refresh SoTA packs, trace provenance, or improve pattern quality. Also triggers on: FPF, bounded contexts, SoTA packs, assurance calculus, decision theory, causal reasoning, temporal reasoning, architecture description, modularity, constraint-governed unfolding, narrative rendering, structural adequacy, cultural evolution, quality gates, lexical discipline, FPF Parts A-I. Not for simple task planning, general philosophy, or Agile unrelated to FPF.
hooks-managementManage hooks and automation for coding agents (Claude Code, Codex CLI, OpenCode). Use when users want to add, list, remove, update, or validate hooks. Triggers on requests like "add a hook", "create a hook that...", "list my hooks", "remove the hook", "validate hooks", or any mention of automating agent behavior with shell commands or plugins.
investigating-repository-historyInvestigate GitHub repository history before risky code changes using git blame/log, GitHub PRs, review comments, squash/rebase/cherry-pick/rename heuristics, and cited evidence. Use when asking why code exists, whether a change is safe, what PR introduced behavior, or before editing API, compatibility, security, concurrency, persistence, migration, or performance-sensitive code.
maintaining-macos-healthHands-on playbook for macOS disk cleanup, dev-machine optimization, and proactive health alerting. Use when the Mac is full or slow, when a process persistently burns CPU, when a kernel panic / watchdog timeout / vm-compressor-space-shortage / Jetsam event happened, when the user asks to free disk space, audit storage, set up disk/memory/CPU alerts, or restore the same monitoring on a new Mac. Built around Mole (`mo` CLI) for safety guards plus a custom LaunchAgent-based alerter for active warnings. Covers Apple Silicon laptops with heavy AI/Docker workloads. Not for general macOS support, hardware diagnostics, networking issues, GUI / window-manager bugs, Time Machine recovery, or broken app installs.
maintaining-windows-healthHands-on playbook for Windows 11 disk cleanup, dev-machine optimization, and proactive health alerting. Use when the PC is full or slow, when a BSOD / Kernel-Power 41 / crash dump / commit-memory pressure happened, when the user asks to free disk space, audit storage, set up disk/memory alerts, or restore the same monitoring on a new PC. Built around native Microsoft-supported tooling (Storage Sense, cleanmgr, DISM, pnputil, vssadmin, wevtutil, powercfg) as the safety floor, a drift-protected HTML cleanup UI, and a Task Scheduler + BurntToast alerter. Covers dev machines with heavy AI/Docker/WSL workloads. Not for general Windows support, hardware diagnostics, GPU/driver troubleshooting, antivirus/malware removal, Windows Update repair, networking, or app-specific performance problems unrelated to disk or memory pressure.