Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/owl-listener/designpowers/usability-testingnpx skills add Owl-Listener/designpowers --skill usability-testinggit clone --depth 1 https://github.com/Owl-Listener/designpowersWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00031 | $0.00661 |
| Opus 5 | $0.00015 | $0.00331 |
| Sonnet 5 | $0.00006 | $0.00132 |
| Haiku 4.5 | $0.00003 | $0.00066 |
Grade A, and why
usability-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 79 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Usability Testing
Testing with real people is how you find out if the design works — not by looking at it, but by watching someone use it.
When to Use
- After design-builder produces a working prototype
- Before declaring a design complete
- When the design-critic flags persona coverage gaps
- When assumptions about user behaviour need evidence
Process
Step 1: Define What You're Testing
Write 3-5 task scenarios that map to the core jobs in the brief. Each task:
- Starts with a realistic trigger — "You just bought a new plant and want to add it to the app"
- Has a clear success condition — "The plant appears in your list with a watering schedule"
- Does not tell the user how — never say "tap the + button"
Step 2: Select Participants
Recruit 5-8 participants. At minimum include:
- 1 person who uses a screen reader
- 1 person over 60
- 1 person who is not a native speaker of the interface language
- 1 person with low tech confidence
Reference inclusive-personas for the ability spectrum.
Step 3: Choose Method
| Method | When to use | Minimum participants |
|---|---|---|
| Moderated think-aloud | New flows, complex interactions | 5 |
| Unmoderated remote | Simple tasks, large sample | 8-12 |
| Guerrilla (hallway) | POC validation, time-constrained | 3-5 |
| Accessibility audit with AT users | After build, before ship | 2-3 |
Step 4: Write the Test Script
- Welcome — explain what you're testing (the design, not them)
- Background — 2-3 questions about their relationship to the problem
- Tasks — present each scenario one at a time, observe silently
- Debrief — "What was hardest?" "What would you change?"
Never help during a task. Silence is data.
Step 5: Analyse Findings
Classify outcomes per task:
- Completed easily — no hesitation, no errors
- Completed with difficulty — hesitation or errors but recovered
- Failed — could not complete
- Completed wrong — thought they succeeded but didn't
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 79 lines · 31 tokens per session scan A 96a3660b5ea2
usability-testing is a skill published in the GitHub repository Owl-Listener/designpowers (240 stars, last pushed 2mo ago), licensed MIT. It adds 31 tokens to every session and 661 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
doubt-driven-review
In-flight adversarial check on a non-trivial decision BEFORE it stands — distinct from post-hoc review of a finished diff. Use on "stress-test this decision", "are we sure about this", "verify before commit", "poke holes in this", when working in unfamiliar code, or before an irreversible step (migration, prod deploy…
release-cut
Cut a new pi-agent-dashboard release: promote ## [Unreleased] in CHANGELOG.md, bump every workspace package.json per SemVer, commit, tag v , and push — triggering the Release workflow that publishes every non-private workspace, builds the Electron artifacts, and creates a GitHub Release. Use on "cut a release"…
ship-it
Worktree-side implementation orchestrator for an OpenSpec change. Idempotent: gates automated scenarios on filesystem reality, owns the red-test fix loop, runs the docker harness with always-teardown, then drives ship-change inline. Escape hatch writes SHIPITBLOCKED.md. Runnable headless. Triggers: "ship it", "build…
faq-mine
Mine docs/faq.md from README.md, docs/.md, and the pi-hermes memory stores. Dispatches @fast subagents per source, dedupes against the existing FAQ, and merges entries in caveman style. Use when asked to "build / regenerate / extend the FAQ", "mine docs into FAQ", "mine hermes memory into FAQ", "surface runtime…
session-to-guideline
Turn a pi session into a Markdown "how-we-did-it" collaboration guideline: reads the session's JSONL transcript and synthesizes a reusable playbook of which prompts worked, what had to be steered, and how to reproduce the result faster. Use when: "document this session", "write up how we did X with the AI", "make a…
scenario-design
Draft real-life test SCENARIOS (not smoke tests) from a change/feature spec. Derives edge-case, performance, frontend-quirk and error-handling scenarios with ISTQB techniques, routes each to a test level, and writes test-plan.md, emitting clarification questions on a spec gap. Use on "design test scenarios", "what…