Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/posit-dev/py-shiny-site/testing-example-appsnpx skills add posit-dev/py-shiny-site --skill testing-example-appsgit clone --depth 1 https://github.com/posit-dev/py-shiny-siteWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00073 | $0.01721 |
| Opus 5 | $0.00036 | $0.00860 |
| Sonnet 5 | $0.00015 | $0.00344 |
| Haiku 4.5 | $0.00007 | $0.00172 |
Grade A, and why
testing-example-apps scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 123 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Testing Example Apps
Overview
Component example apps (components/**/app.py, components/**/app-*.py) are the source
of truth for each page — the shinylive links in index.qmd are generated from them. Smoke
coverage (it loads with no server, JS, or output errors) is centralized: a single
parametrized components/test_examples_smoke.py auto-discovers and smoke-tests every such app,
one test_example_app_smoke[<relative path>] case per app. This skill covers adding
py-shiny Playwright
interaction tests for the primary Core + Express apps via py-shiny's controllers — you
do not write per-component smoke tests.
Everything uses the public shiny package API — shiny.pytest.create_app_fixture,
shiny.playwright.controller, shiny.run.ShinyAppProc. No custom test runner.
Reference implementation: components/layout/accordion/test_accordion.py.
Shared infrastructure (already in place — do not re-create)
pytest.ini(repo root):testpaths = componentsscopes pytest to the site's tests and away from thepy-shiny/submodule.components/conftest.pyprovides:create_example_fixture(HERE / "app-NAME.py")— returns a fixture yielding a runningShinyAppProc. Splits multi-file## file:shinylive apps into a temp dir; launches single-file apps directly.smoke_testfixture — a callablesmoke_test(page, app, *, allow_stderr=(), allow_js=())that navigates, waits for Shiny idle, and asserts no un-allow-listed server stderr, no JS console errors, and zero.shiny-output-error.example_app_paths()/launch_example_app()— power the centralized smoke sweep below; you generally don't call these directly from a component'stest_<name>.py.
components/test_examples_smoke.py— smoke-tests EVERY discovered example app (app.py/app-*.py) as its own parametrized case; you do NOT write per-component smoke tests.components/test_component_pages.py— a static (non-browser) check that every component page (components/<section>/<name>/index.qmd) ships at least one example app (app.py/app-*.py). All discovered pages are enforced; a page may only opt out viaEXEMPT_PAGES(normally empty) with a documented reason. Runs undermake test-components-examples.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 123 lines · 73 tokens per session scan A 3382348638e1
testing-example-apps is a skill published in the GitHub repository posit-dev/py-shiny-site (22 stars, last pushed 19d ago), licensed MIT. It adds 73 tokens to every session and 1,721 once invoked, about $0.0004 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
systematic-debugging
Use when encountering any bug, test failure, or unexpected behavior, before proposing fixes.
brainstorming
You MUST use this before any creative work - creating features, building components, adding functionality, or modifying behavior. Explores user intent, requirements and design before implementation.
chat-pet-sprite-creation
Use when creating or changing VS Code chat pet sprite art, sprite sheets, state animations, eye treatments, Stable/Insiders variants, or pet transitions under src/vs/workbench/contrib/chat/browser/widget/media/chatPet.
cpu-profile-analysis
Analyze V8/Chrome CPU profiles (.cpuprofile) and DevTools trace files (Trace-.json). Use when: profiling performance, investigating slow functions, comparing code paths, finding bottlenecks, analyzing timeToRequest, understanding call trees from sampling profiler data, analyzing layout/paint/rendering, investigating…
agent-host-chat-contributions
Build and review cross-cutting agent-host chat behavior through lifecycle contributions. Use when adding turn lifecycle side effects, prompt or context injection, restored-history transformation, protocol-action observation, or when reviewing changes that add code to AgentSideEffects or AgentService.
auto-perf-optimize
Run agent-driven VS Code performance or memory investigations. Use when asked to launch Code OSS, automate a VS Code scenario, run the Chat memory smoke runner, capture renderer heap snapshots, take workflow screenshots, compare run summaries, or drive a repeatable scenario before heap-snapshot analysis.