Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add fmind/dot --skill python-testinggit clone --depth 1 https://github.com/fmind/dotWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/fmind/dot/python-testing)<a href="https://agentmods.dev/skills/fmind/dot/python-testing"><img src="https://agentmods.dev/badge/skills/fmind/dot/python-testing/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/fmind/dot/python-testing"><img src="https://agentmods.dev/badge/skills/fmind/dot/python-testing.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00019 | $0.01043 |
| Opus 5.5 | $0.00008 | $0.00417 |
| Sonnet 5.5 | $0.00004 | $0.00209 |
| Haiku 4.5 | $0.00002 | $0.00104 |
Grade A, and why
python-testing scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 50 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Python Testing
Prove a change with an honest red-green-refactor cycle: a failing test that detects the missing or broken behavior, then the smallest trustworthy change; quality-assurance owns the broader campaign and systematic-debugging owns failures not yet understood.
For pytest fixture, collection, or assertion maintenance, use pytest mechanics directly. Use the red-green-refactor workflow below when implementing a behavior change; test-only maintenance does not require inventing a production change.
Workflow
- Discover the harness: Read repository instructions, existing tests, task definitions, and nearby patterns; find the smallest command that exercises the target behavior.
- State the contract: Name the production change that would make the test pass and a plausible regression that would make it fail.
- RED: Write one minimal test for one observable behavior; run it and confirm it fails for the missing behavior, not for a syntax, fixture, environment, or setup error.
- GREEN: Implement only enough production code to satisfy that test; run the focused test and read the full output.
- Protect the neighborhood: Run the package or subsystem tests; fix production code when the new behavior breaks a valid existing contract and revisit the spec when contracts conflict.
- REFACTOR: Improve names, structure, duplication, and types only while everything stays green; add no behavior.
- Repeat: Take the next smallest behavior, edge case, or failure path through a new red cycle.
- Prove the regression test: Reuse the observed red result. If implementation preceded the test or its ability to detect the defect remains uncertain, safely exercise the test against the unfixed code in isolation, then confirm green with the fix.
- Qualify proportionately: reuse passing focused and subsystem results while relevant inputs remain unchanged; add affected static checks. Run the full gate only when repository policy or cross-cutting risk requires it. Apply the dirty-tree rule when unrelated work is present.
- Report evidence: summarize the observed red and green outcomes, checks actually run, and remaining limits; do not run additional suites merely to fill the report.
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago Changed 47243fd8de73
- 7d ago Changed c01950a0d220
- 10d ago Changed 3203f1aabf8b
- 13d ago First seen · 50 lines · 19 tokens per session scan A 8441a64a4c38
python-testing is a skill published in the GitHub repository fmind/dot (9 stars, last pushed yesterday), licensed MIT. It adds 19 tokens to every session and 1,043 once invoked, about $0.0001 per session on Opus 5.5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-17.
Other skills, from other repositories
python-stack
Build typed Python projects with uv, Ruff, ty, pytest, Litestar, and Typer. Use for packages, CLIs, web apps, tests, typing, or API verification.
python-script
Write standalone single-file Python scripts using PEP 723 inline metadata and uv. Use when creating a quick CLI script that needs dependencies without a full project.
test-driven-development
Implement an isolated bug fix or behavior change with an honest red-green-refactor cycle. Use when a regression test should fail before the correction, or when tested logic, refactors, or seams must prove correctness.
django-tdd
Use when testing Django applications with pytest - TDD workflow, pytest-django setup, factoryboy and model-bakery fixtures, DRF API testing, mocking and patching, integration tests, or coverage.
python
Python development with ruff, mypy, pytest - TDD and type safety.
python-testing
A Python testing guide covering pytest, test coverage, and test-driven development (TDD), a method of writing a failing test before the code that makes it pass.