sre-architect

sre-architect is an agent for Claude Code from DDS-Solutions/AI-TadPole-OS. It costs 33 tokens per session (1,186 once invoked), scanned A, original, MIT.

A reliability specialist for keeping software available, understandable when it fails, and quick to recover. It covers service targets, failure investigation, incident response, and testing systems under stress.

In plain words
What is it for?
Use it to define reliability targets, choose useful signals such as delay and errors, investigate incidents, improve recovery, and test failure handling before production.
Why use it?
It helps prevent quiet failures, chains of related outages, excessive alerts, and recovery that takes too long. Observability means collecting enough useful evidence to understand why a problem happened.

Agent for Claude Code

Written for Claude Code: a Claude Code subagent (agents/*.md). Also seen: model in frontmatter.

Good fit Use it to define reliability targets, choose useful signals such as delay and errors, investigate incidents, improve recovery, and test failure handling before production.

Compare 6 agents from other repositories ↓
Install with agentmods
npx agentmods add agents/dds-solutions/ai-tadpole-os/sre-architect
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Clone the repo
git clone --depth 1 https://github.com/DDS-Solutions/AI-TadPole-OS

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for sre-architect

README.md
[![agentmods](https://agentmods.dev/badge/agents/dds-solutions/ai-tadpole-os/sre-architect.svg)](https://agentmods.dev/agents/dds-solutions/ai-tadpole-os/sre-architect)
Your own site
<a href="https://agentmods.dev/agents/dds-solutions/ai-tadpole-os/sre-architect"><img src="https://agentmods.dev/badge/agents/dds-solutions/ai-tadpole-os/sre-architect.svg" alt="Measured on agentmods" height="20"></a>
Per session 33 Only the description is in the session, so the agent can decide to use it. The body loads when it is invoked.
When invoked 1,186 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00033 $0.01186
Opus 5 $0.00016 $0.00593
Sonnet 5 $0.00007 $0.00237
Haiku 4.5 $0.00003 $0.00119

Measured 8d ago against content hash ba509e53caec, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-08, from the pricing page.

Security

Grade A, and why

sre-architect scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.agent/agents/sre-architect.md · 71 lines

How it starts

The opening of the file, as written. The whole thing — 71 lines — stays where its author put it; the contents beside it link to each section on GitHub.

[!IMPORTANT] AI Context & Knowledge Heritage

  • Subsystem: Specialist Agent Profiles / sre-architect
  • Architecture: @docs ARCHITECTURE:Documentation
  • Failure Path: "Silent" failures, cascading timeouts, alert fatigue, or recovery times that exceed the business's tolerance.
  • Observability: Traceability via execution/parity_guard.py ([sre_architect])

SRE Architect

Hope is not a strategy. Reliability is a feature. Build for the crash.

🏛️ Philosophy

  • The Error Budget: Reliability is not 100%. We define a tolerable error rate. If the budget is spent, feature velocity stops and stability work begins.
  • Observability > Monitoring: Monitoring tells you that something is broken. Observability tells you why it is broken without needing to deploy new logs.
  • MTTR is the Only Metric: Mean Time To Recovery is the ultimate measure of success. A system that fails is fine; a system that cannot be recovered quickly is a disaster.
  • Anti-Fragility: Use chaos engineering to break the system in staging so it is impossible to break in production.

🛠️ Reliability Frameworks

  • The Golden Signals: Latency, Traffic, Errors, and Saturation.
  • Cascading Failure Prevention: Implement circuit breakers, bulkhead patterns, and exponential backoff with jitter.
  • Sovereign Recovery: Automated health checks $\rightarrow$ Automatic instance replacement $\rightarrow$ Traffic shifting.
  • Stateful Recovery: Database Point-in-Time Recovery (PITR) and cross-region replication.

🧠 Aletheia Reasoning Protocol (Reliability)

1. Generator (The Signal)

  • SLI Definition: "What is the one metric that truly defines if the user is happy? (e.g., 'The /checkout API returns 200 OK within 500ms')."
  • Failure Mode Analysis: "If the Redis cache goes down, does the database collapse under the sudden load? (The Thundering Herd problem)."
  • Capacity Projection: "At 10x current traffic, where is the first bottleneck? CPU, Memory, I/O, or Connection Pool?"

Read the full file on GitHub · 71 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 8d ago First seen · 71 lines · 33 tokens per session scan A ba509e53caec

Subscribe to this mod's changes

sre-architect is an agent published in the GitHub repository DDS-Solutions/AI-TadPole-OS (8 stars, last pushed today), licensed MIT. It adds 33 tokens to every session and 1,186 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other agents, from other repositories

performance-profiler

Secondary reviewer for the performance lens — startup, memory, CPU, rendering, Rust hot paths, AI request efficiency, token optimization, and export performance. Activates only on perf-sensitive changes (hot paths in export/, scraping/, ai/, large lists, SQLite-on-tokio) as a Secondary alongside the domain Primary.

saeedkolivand/ai-job-hunter-app · 68 tokens

webgl-perf-profiler

Cross-cutting GL frame-rate profiler for apps/landing - measures the worst t-segment via a Chrome DevTools performance trace, then applies the landing degradation ladder IN ORDER, stopping at the first rung that passes. Has write access to apply rungs. GL frame-rate only - distinct from performance-profiler (desktop…

saeedkolivand/ai-job-hunter-app · 82 tokens

frontend-developer

Frontend feature lead for cross-cutting frontend work — module architecture, component boundaries, React/TypeScript/CSS/API integration, accessibility, and final quality gates. Detects the project's framework and stack before acting. Use for frontend tasks spanning multiple concerns; for narrow work prefer the focused…

ivklgn/ai-kit · 84 tokens

react-typescript-specialist

Use this agent when you need to develop React components with TypeScript, implement modern React patterns with strict type safety, or refactor existing React code to follow TypeScript best practices. Examples: Context: User needs to create a new React component with proper TypeScript typing. user: 'I need to create a…

PacktPublishing/Agentic-Coding-with-Claude-Code · 242 tokens

migration-specialist

Use this agent when upgrading React versions, migrating between frameworks (CRA to Vite, Pages to App Router), updating major dependencies, running codemods, or handling breaking changes. This agent specializes in safe, incremental migration strategies.

PMDevSolutions/Aurelius · 50 tokens

Frontend Developer

Specialist for TypeScript/React frontend code within a vertical slice. Implements React components that consume auto-generated command and query proxies, following the project's component and styling conventions.

Cratis/AI · 37 tokens