Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/chen3feng/agent-skills/test-layout-evolutionnpx skills add chen3feng/agent-skills --skill test-layout-evolutiongit clone --depth 1 https://github.com/chen3feng/agent-skillsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/chen3feng/agent-skills/test-layout-evolution)<a href="https://agentmods.dev/skills/chen3feng/agent-skills/test-layout-evolution"><img src="https://agentmods.dev/badge/skills/chen3feng/agent-skills/test-layout-evolution.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00035 | $0.01612 |
| Opus 5 | $0.00017 | $0.00806 |
| Sonnet 5 | $0.00007 | $0.00322 |
| Haiku 4.5 | $0.00003 | $0.00161 |
Grade A, and why
test-layout-evolution scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 3d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 154 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Test layout evolution: unit tests vs. legacy test/
When to use
You are about to add the first real unit tests to a repository
that already has a directory called test/ (or tests/), but on
inspection that directory is not unit tests — it drives the
real binary end-to-end against fixture projects, boots a server,
shells out, or otherwise takes seconds per case.
Signal phrases: "there's already a test/ folder, I'll add mine
in there", "let's just add a test_xxx.py next to the existing
integration script", or a PR diff where a 10-ms mocked test
suddenly depends on testdata/, PATH, or network.
Problem
Conflating the two kinds of tests causes recurring pain:
- Speed and feedback loop. Integration suites are slow (real subprocess, real filesystem, real network) and can't run on every save. Unit suites are ~10ms and should run on every save. Put them in one directory and CI either pays the integration cost every time or never runs the units.
- Import / discovery conflicts. The legacy
test/is often not structured as a Python package at all (it's shell scripts- a driver). Dropping
test_*.pynext torun_tests.shconfuses pytest's collection and custom runners.
- a driver). Dropping
- Fixtures mismatch. Integration fixtures tend to live in
testdata/with real files; unit tests prefer in-memory mocks. Mixing them invites a unit test to accidentally open a fixture file and become a slow integration test with a unit test's name. - Dependency creep. Unit tests mock
run_command, network, clocks. Integration tests need those for real. Oneconftest.pytrying to serve both grows monkey-patches that silently break the integration side. - "Just refactor the old tests" is a trap. The integration suite is usually load-bearing for release gating and has accumulated non-obvious dependencies (env vars, working-dir assumptions, testdata). Moving it is a multi-PR project that should not block shipping the first unit test.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 3d ago First seen · 154 lines · 35 tokens per session scan A 210c896345a3
test-layout-evolution is a skill published in the GitHub repository chen3feng/agent-skills (5 stars, last pushed 3mo ago), licensed Apache-2.0. It adds 35 tokens to every session and 1,612 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
run-helix-tests
Submit and monitor .NET MAUI unit tests on Helix infrastructure. Supports running XAML, Resizetizer, Core, Essentials, and other unit test projects on distributed Helix queues.
golang-testing
Production-ready Golang tests — table-driven tests, testify suites and mocks, parallel tests, fuzzing, fixtures, goroutine leak detection with goleak, snapshot testing, code coverage, integration tests, idiomatic test naming. Use when writing or reviewing Go tests, choosing a testing approach, setting up Go test CI…
nunit-testing
Use when writing or modifying tests in NUnit's own test projects, or when making a behavioral change to production code that needs test coverage. Covers test structure, attribute choice, helper visibility, platform guards, and which test projects are real.
ghttp
The ghttp test HTTP server for testing HTTP clients — NewServer/NewTLSServer, AppendHandlers, the Verify assertions (VerifyRequest/VerifyHeader/VerifyHeaderKV/VerifyJSON/VerifyForm/VerifyBasicAuth/VerifyContentType/VerifyBody), RespondWith/RespondWithJSONEncoded(Ptr), CombineHandlers, RouteToHandler for unordered…
detect-flaky-tests
Detects flaky Go tests by analyzing GitHub Actions workflow runs across the last 7 days and all PRs — covering both the run-tests job (unit/integration) and the e2e-test job (gVisor and microVM lanes). For each newly-detected flaky test or infra issue, opens a GitHub issue with full evidence and a draft fix PR. Does…
add-go-test
Write or extend Go unit tests in this repo. Use when the user asks to add or update Go tests.