Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add gabrielmoreira/agent-skills-mirror --skill e2e-setupgit clone --depth 1 https://github.com/gabrielmoreira/agent-skills-mirrorWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/gabrielmoreira/agent-skills-mirror/e2e-setup)<a href="https://agentmods.dev/skills/gabrielmoreira/agent-skills-mirror/e2e-setup"><img src="https://agentmods.dev/badge/skills/gabrielmoreira/agent-skills-mirror/e2e-setup/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/gabrielmoreira/agent-skills-mirror/e2e-setup"><img src="https://agentmods.dev/badge/skills/gabrielmoreira/agent-skills-mirror/e2e-setup.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00101 | $0.00987 |
| Opus 5 | $0.00051 | $0.00494 |
| Sonnet 5 | $0.00020 | $0.00197 |
| Haiku 4.5 | $0.00010 | $0.00099 |
Grade A, and why
e2e-setup scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 69 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Set up an e2e test suite
E2e tests verify the whole running system through the app (browser/API), not one
module. They are the per-PR gate. Pairs with dev-local-setup (a reproducible
local stack) and verifier-setup (which scaffolds the repo's /verify skill —
the verify→ship loop this gate feeds into).
Where it lives
- Unit/integration tests stay inside each app/package — they own one module.
- System e2e is a dedicated top-level package (e.g.
e2e/) — it spans all apps, so it belongs to none. Add it to the workspace if a monorepo.
The recipe
- Stand the app up reproducibly — see
dev-local-setup. The e2e suite never boots the app itself; it runs against the already-running stack. That stack can be local (dev-local-setup) or an isolated cloud box (crabbox-setup) — same specs, run against either. For parallel agents use the cloud box (one laptop can't host concurrent stacks). - Pick the framework that fits (Playwright for browser; your HTTP client for API). Turn on video + trace — the recording is the proof, and it's gitignored output.
- Explore the flow live first (don't guess selectors), then crystallize it into a committed spec.
- Keep the gate small: a handful of critical journeys, deterministic. Each new feature PR adds its spec — the suite compounds.
Practices that make e2e trustworthy
- Real flow, not bypass. Drive the genuine path. For email codes / OTP, read the real code from a local mail server (Mailpit / Inbucket / MailHog) — never hardcode a fixed test code. That's what makes it a test, not a rehearsal.
- Verify auth ITSELF once; bypass it everywhere else. A dedicated signup/login spec proves auth works. Every other spec shouldn't re-pay the login tax — build a session helper that mints an authed state once (real flow → saved storage state, or a service-role/token mint) and load it.
- Layered assertions: client → server → product. Don't stop at "the UI changed." Confirm the server agrees (token validates / row/state is right) AND the user-visible outcome (e.g. plan upgraded and credits granted).
- Stable selectors. Prefer role/label/text; add a small
data-testidin the component when there's no good handle — never a brittle CSS path. - Fresh data per run. Unique emails/ids so reruns don't collide; mind rate limits (auth email, etc.).
- Commit specs + helpers, never
test-results/(generated output).
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 69 lines · 101 tokens per session scan A 1e7054495cb4
e2e-setup is a skill published in the GitHub repository gabrielmoreira/agent-skills-mirror (17 stars, last pushed yesterday), licensed MIT. It adds 101 tokens to every session and 987 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-03.
Other skills, from other repositories
reproduce-bug
Reproduce a reported bug in googleapis/mcp-toolbox and decide whether it is real, delivering an evidence-backed verdict: confirmed, already fixed, misconfiguration, client-side, works as intended, not reproducible, or blocked. Use whenever a maintainer asks you to reproduce, verify, confirm, or investigate a bug…
verify-bug
Post-merge UAT verification workflow. Walks JIRA reproduce steps, performs comparative audits (Before/After), attaches evidence to JIRA, and transitions status on PASS.
reality-verification
This skill should be used when the user asks to "verify a fix", "reproduce failure", "diagnose issue", "check BEFORE/AFTER state", "VF task", "reality check", "check test quality", "mock-only tests", or needs guidance on verifying fixes by reproducing failures before and after implementation, or detecting mock-heavy…
testing-blocks
Use this when you have made AEM Edge Delivery Services code changes to blocks, scripts, or styles and need to validate them before opening a pull request. Covers unit testing for utilities and logic, browser testing with Playwright, linting, and guidance on what to test and how.
test-electron-app
Drive the real running PostHog Electron app (live tRPC, workspace-server, real data) over CDP with agent-browser. Connect to the running app on port 9222, test desktop changes against a local Django stack, snapshot the accessibility tree, inspect network requests, and screenshot only when explicitly asked. Use when…
qa
Read-only automated QA sweep of a deployed stardust site on AEM Edge Delivery Services — validates routing, content fidelity vs the source capture, template conformance, rendered integrity (geometry, JS errors, broken images), visual regression vs baselines, metadata/SEO/JSON-LD, link integrity, accessibility (axe)…