naive-test

A command for testing a live application as a first-time user who does not know how it was built. It explores the interface and produces a dated findings report.

In plain words
What is it for?
It is for checking an app’s real browser flows before deployment, including sign-in, forms, buttons, links, and other visible interactions. It requires the configured app and browser automation tools.
Why use it?
It finds confusing behavior, broken expectations, and accessibility or usability problems that scripted tests may miss. The test is based on what a user sees and does, not on the source code.

Command

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add commands/techgardencode/naive-user/naive-test
Clone the repo
git clone --depth 1 https://github.com/TechGardenCode/naive-user
Per session 18 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 218 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00018 $0.00218
Opus 5 $0.00009 $0.00109
Sonnet 5 $0.00004 $0.00044
Haiku 4.5 $0.00002 $0.00022

Measured yesterday against content hash 301729b9181f, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

naive-test scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

commands/naive-test.md · 18 lines

What it actually says

Run the naive-user skill against the app named in $ARGUMENTS, or, if none is given, the app in naive-user.config.json.

Invoke the skill and follow it exactly: read naive-user.config.json, make sure the app is running, sign in via the configured auth steps, then explore the live UI as an uninformed first-time user, source-blind, producing an updated mental model and a dated findings report under qa/naive-user/<app>/.

This needs the Playwright MCP server (browser tools). See the project README if the browser_* tools are not available. Run it on demand while developing, or dispatch it as a subagent to run in parallel. It complements scripted end-to-end tests: those assert known flows, this discovers unknown gaps.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 18 lines · 18 tokens per session scan A 301729b9181f

Subscribe to this mod's changes

naive-test is a command published in the GitHub repository TechGardenCode/naive-user (1 stars, last pushed 1mo ago), licensed MIT. It adds 18 tokens to every session and 218 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.