testing-python

A guide to writing Python tests with pytest, a tool for checking that Python code behaves as expected.

In plain words
What is it for?
Use it to design tests, fixtures, parameterized cases, mocks, asynchronous tests, and checks for test failures or missing coverage.
Why use it?
It helps avoid tests that are flaky, unclear, or tied too closely to implementation details.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/ai-riksarkivet/ra-mcp/testing-python
Any agent
npx skills add AI-Riksarkivet/ra-mcp --skill testing-python
Clone the repo
git clone --depth 1 https://github.com/AI-Riksarkivet/ra-mcp

Made for: Claude Code, Codex.

Per session 48 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 3,141 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00048 $0.03141
Opus 5 $0.00024 $0.01571
Sonnet 5 $0.00010 $0.00628
Haiku 4.5 $0.00005 $0.00314

Measured 2d ago against content hash 677deb39d3eb, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

testing-python scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

with patch("myapp.client.requests.get", return_value=mock_response) as mock_get:
packages/mcps/viewer-mcp/.claude/skills/testing-python/SKILL.md · 490 lines

How it starts

The opening of the file, as written. The whole thing — 490 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Writing Effective Python Tests

Core Principles

  • Every test should be atomic, self-contained, and test one behavior
  • Tests should be deterministic — no flaky results from shared state, timing, or network
  • Test the contract (inputs/outputs/errors), not the implementation details
  • Prefer real implementations over mocks when practical

Test Structure

One behavior per test

Each test should verify a single behavior. The test name should tell you what's broken when it fails. Multiple assertions are fine when they all verify the same behavior.

# Good: name tells you what's broken
def test_user_creation_sets_defaults():
    user = User(name="Alice")
    assert user.role == "member"
    assert user.id is not None
    assert user.created_at is not None

# Bad: if this fails, what behavior is broken?
def test_user():
    user = User(name="Alice")
    assert user.role == "member"
    user.promote()
    assert user.role == "admin"
    assert user.can_delete_others()

Arrange-Act-Assert (AAA)

def test_transfer_reduces_sender_balance():
    # Arrange
    sender = Account(balance=100)
    receiver = Account(balance=50)

    # Act
    transfer(sender, receiver, amount=30)

    # Assert
    assert sender.balance == 70

Naming: test_<subject>_<scenario>

# Good — describes subject and scenario
def test_login_fails_with_invalid_password(): ...
def test_parse_csv_skips_empty_rows(): ...
def test_cache_expires_after_ttl(): ...

# Bad — too vague
def test_login(): ...
def test_1(): ...
def test_it_works(): ...

Parameterization for variations of the same concept

import pytest

@pytest.mark.parametrize("input_val,expected", [
    pytest.param("hello", "HELLO", id="lowercase"),
    pytest.param("World", "WORLD", id="mixed-case"),
    pytest.param("", "", id="empty-string"),
    pytest.param("123", "123", id="digits-unchanged"),
])
def test_uppercase_conversion(input_val, expected):
    assert input_val.upper() == expected

Read the full file on GitHub · 490 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 490 lines · 48 tokens per session scan A 677deb39d3eb

Subscribe to this mod's changes

testing-python is a skill published in the GitHub repository AI-Riksarkivet/ra-mcp (18 stars, last pushed 6d ago), licensed Apache-2.0. It adds 48 tokens to every session and 3,141 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

agent-code-analyzer

Agent skill for code-analyzer - invoke with $agent-code-analyzer.

ruvnet/ruflo · 19 tokens

foundry-config-setup

Resolve missing setup caused by a hardcoded Foundry project endpoint or model in a sample. Use when a sample fails because it uses a placeholder/hardcoded projectendpoint (for example "https://your-project.services.ai.azure.com") or a hardcoded model instead of reading them from the environment.

microsoft/agent-framework · 65 tokens

haiku

When writing a haiku for this bot, follow these conventions.

agno-agi/agno · 0 tokens

deploy-docker-compose

Run the Omnigent server as a Docker compose stack (server + Postgres) on any Docker host — your laptop, a VPS, EC2 by hand, or as the base layer of any container-platform deploy. Invoke when the user wants to build the image, bring up the compose stack, debug the stack on a host they already have, or extend the stack…

omnigent-ai/omnigent · 84 tokens

azure-mgmt-botservice-dotnet

Azure Resource Manager SDK for Bot Service in .NET. Management plane operations for creating and managing Azure Bot resources, channels (Teams, DirectLine, Slack), and connection settings. Triggers: "Bot Service", "BotResource", "Azure Bot", "DirectLine channel", "Teams channel", "bot management .NET", "create bot".

microsoft/skills · 78 tokens

dogfood

Systematically explore and test a mobile app on iOS/Android with agent-device to find bugs, UX issues, and other problems. Use when asked to dogfood, QA, exploratory test, find issues, bug hunt, or test this app on mobile.

callstack/agent-device · 55 tokens