Use when profiling a mobile app, authoring or sizing mobile TCs, orchestrating a mobile suite run, executing a mobile TC, generating a manual guide, or producing a mobile run report. Mobile-app testing for all manual-qa agents — native iOS/Android via Appium MCP and the Mobitru device farm, and PWA/hybrid via…
Performance, networking, console, and JavaScript analysis. Checks network issues, console errors, and JS problems; approximates load/resource health where full Core Web Vitals aren't available.
Responsive-layout analysis on a real browser — viewport configuration, touch-target sizing, and mobile-layout issues via a resized Playwright viewport.
Use when the user provides an Excel file of test cases or requirements, or asks to import a spreadsheet. Reads .xlsx/.xls spreadsheets (test cases, checklists, requirement matrices) into Markdown tables so an agent can ingest them.
Use when a hypothesis closes, an experiment concludes with a metric movement or a killed assumption, a vendor / data source / external service behaves differently than its docs implied, or a non-obvious product decision needs its rationale preserved — even if the user never says the word 'learning'. Captures a problem…
Use when a hypothesis names no ratified outcome and so cannot be promoted, a quarterly bet is being set, or someone asks what number a bet is actually trying to move. Drafts, stress-tests, and records ratification of outcome anchors in docs/discovery/outcomes.md — measurable customer-behavior metrics with a dated…
Use when personas come up — 'define the personas', 'who are our users', 'make a persona card for X', 'which persona owns this journey' — or proactively whenever a journey, BDD scenario, or hypothesis names an actor that has no card yet or spells one inconsistently. Creates and maintains the canonical persona cards…
Use when starting a work session, when the PO is unsure what to do next, or when promotion, gates, blockers, what's-stuck, where-am-I, or am-I-ready-for-review come up — even without the word 'status'. Reports the whole discovery pipeline as one navigable dashboard — where every hypothesis stands against the promotion…
Use when you're trying to make a hard call, sharpen fuzzy language, interrogate a fresh Hypothesis before promotion, or pressure-test a plan or new initiative — even if you don't say 'grill'. A Socratic, one-question-at-a-time interview that walks the decision tree branch by branch, taking a position on each question…
Use when a raw ask arrives — a feature request, stakeholder ask, support theme, or 'someone wants X' — before anything becomes a Problem or Hypothesis. Verdicts every item Act Now / Plan Next / Collect More Signal / Decline-or-Defer, writes one batch record, and mints accepted in-scope items as Problems carrying…
Use when journeys and the backlog need reconciling — 'turn my journeys into hypotheses', 'map the journeys to the backlog', 'what do my journeys cover', 'convergence pass', 'where are the gaps in the backlog' — or proactively whenever new journey files exist that no hypothesis or epic references. Runs a convergence…
Use when mapping opportunities under an outcome, asking where a hypothesis hangs, spotting solution-shaped 'problems', or when a new Problem or Hypothesis has no parent. Maintains the opportunity–solution tree over existing Discovery artifacts by adding only nodetype / parent frontmatter to existing files and…
Use when deciding what to build next, when validated hypotheses outnumber the team's appetite, when a ranking feels stale, or when someone asks which bet comes first. Ranks the incubating and promotion-ready bets against the active prioritization framework — RICE by default, with WSJF and ICE as config options …
Use when a stakeholder conversation, customer meeting, or user session is being planned or has just happened — even if the user only says 'I'm meeting them Thursday' or pastes raw notes. Runs two modes — PREPARE aggregates every open question, untested critical assumption, and unresolved escalation across the…
Use when the user says 'file a bug', 'comment on JIRA-123', 'write up a decision page', or authors any Atlassian content. Creates well-formatted Jira issues/comments and Confluence pages on both Cloud and Server/Data Center, with mandatory post-creation re-fetch plus repair.
Use when a scope of test cases (or a described backlog, before any cases even exist as files) needs sizing or a cost/time estimate BEFORE automation work starts — presales scoping, a proposal, "how long/much to automate these N cases", "size these cases S/M/L/XL", estimating the framework/CI/foundation work an…
Use when the user asks 'what did this cost', cost per session/role/test case, which role or sub-agent burned the most, tool-call/skill/time breakdowns, 'before vs after' cost comparisons, or wants to audit AI spend over time. Measures the token/cost/time efficiency of AI coding-agent work — per session, per role, per…
Use when the user asks to 'seed the project', 'onboard this repo', 'generate project config', 'create AGENTS.md', or after the scout has explored the codebase. Generates AGENTS.md and .agents/ configuration files for a new project.
Use when a test case needs to become a green, framework-resident test — the build slot of the test-automation pipeline, from case to green test. Six-phase loop (Absorb → Investigate → Automate → Execute → Debug → Handoff), the 12 Hard Rules, the coverage declaration, the Run Report. Orchestration…
Optional always-on usage telemetry for agent teams — hooks capture every session's tokens, cost, time, activity and named case ids into a git-committed ledger (.agents/telemetry/automation/), covering Claude Code, Copilot CLI AND the VS Code Copilot sidebar, so the data survives transcript expiry and accumulates…
Use when the user asks to 'research trends', 'analyze a topic', 'fact-check this', 'verify claims', 'what's the state of X', or hands you a document to vet. A disk-first, checkpointed research workflow with three modes — trend research, topic analysis, and fact-checking.