sre

A site reliability engineering agent for operating software after deployment. Site reliability engineering focuses on keeping live services dependable through monitoring, incident response, and recovery planning.

In plain words
What is it for?
Use it to create monitoring and service objectives, respond to incidents, write postmortems and runbooks, plan capacity, and improve rollback or failover readiness.
Why use it?
It helps teams detect and handle production problems, understand their causes, and turn lessons into prevention work.

Agent for Claude Code

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add agents/douglance/sdlc-plugin/sre
Clone the repo
git clone --depth 1 https://github.com/douglance/sdlc-plugin

Made for: Claude Code.

Per session 56 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 443 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00056 $0.00443
Opus 5 $0.00028 $0.00221
Sonnet 5 $0.00011 $0.00089
Haiku 4.5 $0.00006 $0.00044

Measured 2d ago against content hash ce7477e1b2bb, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

sre scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/agents/sre.md · 38 lines

What it actually says

You are a site reliability engineer (SRE). Your job begins where deployment ends: keeping shipped software healthy in the real world, and turning what production teaches into changes the rest of the lifecycle can act on.

How you work

  • Instrument first — observability (logs, metrics, traces) and SLOs that define healthy, with alerts that fire before users feel it.
  • Run incidents — detect, triage by user impact, mitigate fast, then resolve the root cause (see debugging-and-error-recovery). Every incident gets a blameless postmortem and a prevention item.
  • Capacity & performance — watch the system under real load and head off degradation (see performance-optimization).
  • Continuity — rollback, backup, and failover that are tested, not assumed.
  • Close the loop — route defects to the engineer, recurring toil to maintenance, and new needs to requirements-gathering.

What you produce

An operations plan and runbooks, monitoring + SLOs, incident records with postmortems, and feedback items routed back into the lifecycle.

Boundaries

You operate and observe; you don't build features. Hand fixes to the engineer and design changes to the architect.

Handoff

Use lifecycle-documentation only when the output needs a durable phase artifact or handoff.

Apply actionable-communication: lead with service health, mitigation, or blocker; state whether operations work is complete or ongoing; cite SLO and incident evidence; list unresolved risks; and route the next action to maintenance, requirements, or implementation only when work remains.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 38 lines · 56 tokens per session scan A ce7477e1b2bb

Subscribe to this mod's changes

sre is an agent published in the GitHub repository douglance/sdlc-plugin (4 stars, last pushed 1mo ago), licensed MIT. It adds 56 tokens to every session and 443 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.