Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add liqiongyu/lenny_skills_plus --skill usability-testinggit clone --depth 1 https://github.com/liqiongyu/lenny_skills_plusWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/liqiongyu/lenny_skills_plus/usability-testing)<a href="https://agentmods.dev/skills/liqiongyu/lenny_skills_plus/usability-testing"><img src="https://agentmods.dev/badge/skills/liqiongyu/lenny_skills_plus/usability-testing/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/liqiongyu/lenny_skills_plus/usability-testing"><img src="https://agentmods.dev/badge/skills/liqiongyu/lenny_skills_plus/usability-testing.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00025 | $0.01976 |
| Opus 5 | $0.00013 | $0.00988 |
| Sonnet 5 | $0.00005 | $0.00395 |
| Haiku 4.5 | $0.00003 | $0.00198 |
Grade A, and why
usability-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 138 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Usability Testing
Scope
Covers
- Designing task-based usability studies tied to a specific product decision
- Testing live flows, prototypes, and “faked” implementations (fake door, Wizard of Oz)
- Running moderated sessions (remote or in-person) and capturing high-quality evidence
- Turning findings into a prioritized fix list (including high-ROI microcopy/CTA improvements)
When to use
- “Create a usability test plan and script for .”
- “We need to test a prototype with 5–8 users next week.”
- “Validate a value proposition before building (fake door / Wizard of Oz).”
- “Help me synthesize usability findings into a prioritized backlog.”
When NOT to use
- You need statistically reliable estimates or causal impact (use analytics/experimentation)
- You need open-ended discovery (“what problems do users have?”) without a specific flow to evaluate (use
conducting-user-interviews) - You need a design critique or heuristic review without live user sessions (use
running-design-reviews) - You need to write specs or design docs for a feature, not test an existing flow (use
writing-specs-designs) - You need to apply behavioral/persuasion design patterns to a flow (use
behavioral-product-design); this skill evaluates usability, not designs behavioral nudges - You’re working with high-risk populations or sensitive topics (medical, legal, minors) without appropriate approvals/training
- You don’t have a concrete scenario/flow to evaluate (clarify the decision first)
Inputs
Minimum required
- Product + target user segment (who, context of use)
- The decision this test should inform (what will change) + timeline
- What you’re testing (flow/feature) + prototype/build link (or “recommend stimulus”)
- Platform + environment (web/mobile/desktop; remote/in-person)
- Constraints: session type, number of participants, incentives, recording policy, privacy constraints
Missing-info strategy
- Ask up to 5 questions from references/INTAKE.md.
- If still unknown, proceed with explicit assumptions and list Open questions that would change the plan.
What ships with it
13 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
- eval/eval_config.json 572 B
- eval/SHOWCASE.md 5.8 KB
- eval/with_skill.md 30 KB
- eval/without_skill.md 29 KB
- README.md 992 B
- references/CHECKLISTS.md 2.2 KB
- references/EXAMPLES.md 1.9 KB
- references/INTAKE.md 1.6 KB
- references/RUBRIC.md 5.2 KB
- references/SOURCE_SUMMARY.md 1.6 KB
- references/TEMPLATES.md 3.2 KB
- references/WORKFLOW.md 3.5 KB
- skillpack.json 375 B
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 138 lines · 25 tokens per session scan A f50fb736102d
usability-testing is a skill published in the GitHub repository liqiongyu/lenny_skills_plus (52 stars, last pushed 3mo ago), licensed Apache-2.0. It adds 25 tokens to every session and 1,976 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
plan
Use when a request needs shaping before any code is written — a rough or vague prompt to sharpen, an ambiguous idea to design, or a clear-enough task to decompose. One chain-starter that amplifies the prompt, designs the approach, and decomposes it into a batched task file, skipping whichever phases the request…
decision-heuristics
A set of heuristics for making difficult personal decisions such as changing jobs, buying a home, moving, forming a partnership, or getting married. It is intended for major choices, not everyday decisions.
screen-detox
A habit-change framework for reducing non-work screen use such as social media, short videos, games, and endless scrolling.
monkey-mind-meditation
A meditation and reflection guide for a busy or anxious mind. It treats recurring thoughts as something to observe and examine, while noting that acute serious mental-health problems need professional care.
happiness-skill
A Chinese-language guide to happiness based on reducing unmet wants, focusing on the present, and treating happiness as a trainable skill.
peer-selection
A guide for thinking about which friends, partners, and social circles to keep close, based on shared values and their influence on your life.