Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/onsager-ai/dev-skills/web-testingnpx skills add onsager-ai/dev-skills --skill web-testinggit clone --depth 1 https://github.com/onsager-ai/dev-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/onsager-ai/dev-skills/web-testing)<a href="https://agentmods.dev/skills/onsager-ai/dev-skills/web-testing"><img src="https://agentmods.dev/badge/skills/onsager-ai/dev-skills/web-testing.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00119 | $0.01162 |
| Opus 5 | $0.00060 | $0.00581 |
| Sonnet 5 | $0.00024 | $0.00232 |
| Haiku 4.5 | $0.00012 | $0.00116 |
Grade A, and why
web-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 78 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Web Testing Protocol (L2)
Exploratory, AI-driven validation of dashboard UI changes — not regression testing. Regression coverage is L1's job (smoke + e2e suites under the dashboard's test directories). L2 catches things L1 misses: layout bugs, mobile regressions, interaction flows that only fail in a real browser.
Adopting in another repo: the procedure (read diff → map to routes → run agent-browser at two viewports → screenshot → verdict) is repo-agnostic. The route table, test paths, and verdict schema below are examples from
onsager-ai/onsager. Fork the skill and replace those concrete bits for your own dashboard.
When to invoke
- A PR touches
apps/dashboard/** - L1 e2e fails and you need to know if it's a real regression, flaky, or env
- Someone says "validate the UI" / "dogfood this change"
The app under test
The CI pipeline builds crates/stiglab/deploy/Dockerfile — a single image bundling the Rust backends (stiglab + synodic) and the prebuilt dashboard SPA. It listens on http://localhost:3000.
Primary routes:
| Route | Page | Heading |
|---|---|---|
/ |
Factory overview | Factory |
/sessions |
Sessions list | Sessions |
/sessions/:id |
Session detail | — (dynamic) |
/nodes |
Nodes list | Nodes |
/artifacts |
Artifacts list | Artifacts |
/spine |
Event spine viewer | — (dynamic) |
/governance |
Governance | Governance |
/settings |
Settings + credentials | Settings |
Viewports (always test both)
- Desktop:
agent-browser set viewport 1280 720 - Mobile:
agent-browser set viewport 375 812
Mobile matters — the dashboard ships with a responsive layout (see the md: breakpoints throughout). Horizontal overflow and hidden nav are the top-two regression classes.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 78 lines · 119 tokens per session scan A ca757f9be1c5
web-testing is a skill published in the GitHub repository onsager-ai/dev-skills (5 stars, last pushed 22d ago), licensed MIT. It adds 119 tokens to every session and 1,162 once invoked, about $0.0006 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
playwright-cli
Automate browser interactions, test web pages and work with Playwright tests.
use-agent-browser-for-airi
Test AIRI display-model imports with agent-browser across stage-tamagotchi Electron, stage-web, and stage-pocket mobile web layouts. Use when uploading and verifying contributor-supplied Live2D ZIP, VRM, or MMD ZIP/PMX/PMD files through AIRI's model selector, including onboarding bypass, format-specific import…
playwright-validation
Use when validating UI changes in a branch require Playwright E2E testing. Reviews branch changes, validates UI with Playwright MCP, and adds missing test cases.
feature-walkthrough
Autonomously test the running binance-trading-bot app in a real browser. The agent drives the browser itself via the Playwright MCP - logs in, looks at each screen, and works through every feature end-to-end like a real operator, finding and fixing bugs. Use when asked to test the app, smoke-test or walk through the…
playwright-screen-recording
Record browser test videos with Playwright for PR review and bug fix verification.
browser-smoke-review
Use browser automation to review docs pages, preview URLs, rendered output, or web-facing fallow surfaces. Use when the user wants a screenshot-based review, browser smoke test, docs site check, or preview deployment inspection.