backend-python

A set of rules for building Python web backends with FastAPI, Pydantic, and vLLM. It describes where routes, shared dependencies, errors, schemas, and engine-specific code should live.

In plain words
What is it for?
Use it when adding or changing FastAPI endpoints, request and response schemas, error handling, configuration, or vLLM-backed inference code.
Why use it?
It gives the project a consistent structure and keeps HTTP handling, business coordination, and model-engine logic separate. It also protects endpoint contracts from accidental changes.

Cursor rule for Cursor

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add rules/coeusyk/inference-x/backend-python
Clone the repo
git clone --depth 1 https://github.com/coeusyk/inference-x

Made for: Cursor.

Per session 424 This file is loaded in full into every session.
When invoked 424 The same file — it is already loaded in full.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00424 $0.00424
Opus 5 $0.00212 $0.00212
Sonnet 5 $0.00085 $0.00085
Haiku 4.5 $0.00042 $0.00042

Measured yesterday against content hash 256822926a3e, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

backend-python scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.cursor/rules/backend-python.mdc · 53 lines

What it actually says

Backend Python rules

Stack assumptions

  • Python 3.11+
  • FastAPI for HTTP APIs
  • Pydantic for schemas
  • vLLM as the first engine implementation
  • Development on Windows via WSL2 Ubuntu

Code style

  • Use type hints on all public functions, methods, and class attributes where practical.
  • Prefer small classes and focused functions.
  • Keep route handlers limited to validation, dependency resolution, and response formatting.
  • Put orchestration in services and engine-specific logic in engine modules.
  • Prefer explicit imports over wildcard imports.
  • Raise domain-specific exceptions or HTTP errors deliberately; do not swallow exceptions.

FastAPI conventions

  • Define routes under src/inference_x/api/routes/.
  • Define shared dependencies in src/inference_x/api/deps.py.
  • Define error mapping in src/inference_x/api/errors.py.
  • Keep endpoint contracts stable once introduced.
  • Prefer versioned endpoints under /v1.

Schema conventions

  • Put request and response models in src/inference_x/schemas/.
  • Separate transport schemas from engine internals when complexity grows.
  • Keep field names aligned with OpenAI-compatible formats where intended.
  • Validate config-derived values before use.

Engine conventions

  • Define a base engine interface in src/inference_x/engines/base.py.
  • Implement vLLM-specific behavior in src/inference_x/engines/vllm_engine.py.
  • Keep prompt formatting and request adaptation outside route handlers.
  • Expose health-check capability from each engine implementation.

Testing conventions

  • Unit tests go in tests/unit/.
  • Integration tests go in tests/integration/.
  • Contract tests go in tests/contract/.
  • Mock vLLM where possible in unit tests; reserve real-engine execution for explicit integration tests.

Operational rules

  • Read model and server settings from YAML or environment-backed settings.
  • Do not hardcode model identifiers, ports, file paths, or GPU values in core logic.
  • Log structured events for engine startup, request handling, latency, and failures.
Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 53 lines · 424 tokens per session scan A 256822926a3e

Subscribe to this mod's changes

backend-python is a cursor rule published in the GitHub repository coeusyk/inference-x (2 stars, last pushed 20d ago), licensed MIT. It adds 424 tokens to every session, about $0.0021 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.