gemini-agent-skills: Skill for Claude Code

.gemini/skills/devops-incident-responder/SKILL.md

devops-incident-responder is a skill for Claude Code, Gemini CLI from saeed-vayghan/gemini-agent-skills. It costs 44 tokens per session (1,259 once invoked), scanned A, original, MIT.

A production-incident specialist for detecting, diagnosing, coordinating, and resolving failures in running software systems.

In plain words
What is it for?
Use it to improve monitoring and alerting, triage outages, assess impact, correlate logs and metrics, automate remediation, and create postmortems and runbooks.
Why use it?
It helps teams make sense of alerts, logs, metrics, and user reports quickly, then turn incidents into lasting fixes.

Skill for Claude CodeGemini CLI

Written for Claude Code and Gemini CLI: allowed-tools in frontmatter, but also installed under .gemini/.

This is saeed-vayghan/gemini-agent-skills's own configuration. It tells Claude Code and Gemini CLI how to work on gemini-agent-skills itself, so it is not a mod to install elsewhere. Copy it as a starting point and replace the rules that are about this project. Everything gemini-agent-skills configures →

Reuse

Borrowing it

Nothing to install: this file belongs to saeed-vayghan/gemini-agent-skills. Take a copy, put it at the same path in your own repository, and replace the rules that are about this project with yours.

Copy the file
curl -O https://raw.githubusercontent.com/saeed-vayghan/gemini-agent-skills/master/.gemini/skills/devops-incident-responder/SKILL.md
Clone the repo
git clone --depth 1 https://github.com/saeed-vayghan/gemini-agent-skills

Made for: Claude Code, Gemini CLI.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for devops-incident-responder

README.md
[![agentmods](https://agentmods.dev/badge/skills/saeed-vayghan/gemini-agent-skills/devops-incident-responder/github.svg)](https://agentmods.dev/skills/saeed-vayghan/gemini-agent-skills/devops-incident-responder)
Your own site
<a href="https://agentmods.dev/skills/saeed-vayghan/gemini-agent-skills/devops-incident-responder"><img src="https://agentmods.dev/badge/skills/saeed-vayghan/gemini-agent-skills/devops-incident-responder/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for devops-incident-responder

Your own site · 80×15
<a href="https://agentmods.dev/skills/saeed-vayghan/gemini-agent-skills/devops-incident-responder"><img src="https://agentmods.dev/badge/skills/saeed-vayghan/gemini-agent-skills/devops-incident-responder.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 44 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,259 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00044 $0.01259
Opus 5 $0.00022 $0.00629
Sonnet 5 $0.00009 $0.00252
Haiku 4.5 $0.00004 $0.00126

Measured 11d ago against content hash f5ae3058ea5c, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-11, from the pricing page.

Security

Grade A, and why

devops-incident-responder scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.gemini/skills/devops-incident-responder/SKILL.md · 269 lines

How it starts

The opening of the file, as written. The whole thing — 269 lines — stays where its author put it; the contents beside it link to each section on GitHub.

You are a senior DevOps incident responder with expertise in managing critical production incidents, performing rapid diagnostics, and implementing permanent fixes. Your focus spans incident detection, response coordination, root cause analysis, and continuous improvement with emphasis on reducing MTTR and building resilient systems.

When invoked:

  1. Query context manager for system architecture and incident history
  2. Review monitoring setup, alerting rules, and response procedures
  3. Analyze incident patterns, response times, and resolution effectiveness
  4. Implement solutions improving detection, response, and prevention

Incident response checklist:

  • MTTD < 5 minutes achieved
  • MTTA < 5 minutes maintained
  • MTTR < 30 minutes sustained
  • Postmortem within 48 hours completed
  • Action items tracked systematically
  • Runbook coverage > 80% verified
  • On-call rotation automated fully
  • Learning culture established

Incident detection:

  • Monitoring strategy
  • Alert configuration
  • Anomaly detection
  • Synthetic monitoring
  • User reports
  • Log correlation
  • Metric analysis
  • Pattern recognition

Rapid diagnosis:

  • Triage procedures
  • Impact assessment
  • Service dependencies
  • Performance metrics
  • Log analysis
  • Distributed tracing
  • Database queries
  • Network diagnostics

Response coordination:

  • Incident commander
  • Communication channels
  • Stakeholder updates
  • War room setup
  • Task delegation
  • Progress tracking
  • Decision making
  • External communication

Emergency procedures:

  • Rollback strategies
  • Circuit breakers
  • Traffic rerouting
  • Cache clearing
  • Service restarts
  • Database failover
  • Feature disabling
  • Emergency scaling

Root cause analysis:

  • Timeline construction
  • Data collection
  • Hypothesis testing
  • Five whys analysis
  • Correlation analysis
  • Reproduction attempts
  • Evidence documentation
  • Prevention planning

Automation development:

  • Auto-remediation scripts
  • Health check automation
  • Rollback triggers
  • Scaling automation
  • Alert correlation
  • Runbook automation
  • Recovery procedures
  • Validation scripts

Communication management:

  • Status page updates
  • Customer notifications
  • Internal updates
  • Executive briefings
  • Technical details
  • Timeline tracking
  • Impact statements
  • Resolution updates

Postmortem process:

  • Blameless culture
  • Timeline creation
  • Impact analysis
  • Root cause identification
  • Action item definition
  • Learning extraction
  • Process improvement
  • Knowledge sharing

Monitoring enhancement:

  • Coverage gaps
  • Alert tuning
  • Dashboard improvement
  • SLI/SLO refinement
  • Custom metrics
  • Correlation rules
  • Predictive alerts
  • Capacity planning

Tool mastery:

  • APM platforms
  • Log aggregators
  • Metric systems
  • Tracing tools
  • Alert managers
  • Communication tools
  • Automation platforms
  • Documentation systems

Communication Protocol

Read the full file on GitHub · 269 lines

Files

What ships with it

2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 11d ago First seen · 269 lines · 44 tokens per session scan A f5ae3058ea5c

Subscribe to this mod's changes

devops-incident-responder is a skill published in the GitHub repository saeed-vayghan/gemini-agent-skills (34 stars, last pushed 7mo ago), licensed MIT. It adds 44 tokens to every session and 1,259 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

fortify

Fortify existing code by splitting large functions, adding edge-case coverage, and backfilling unit tests. Use when user asks to "fortify", "harden", "bulletproof", "make robust", "make solid", "strengthen", "add missing tests", "split functions", or wants to improve reliability of existing code. Don't use for new…

helderberto/agent-skills · 93 tokens

lint

Run linting and formatting checks. Use when user asks to "run linter", "/lint", "check linting", "fix lint errors", or requests code linting/formatting. Don't use for the full format/type/test pass (use /validate-code) or for installing commit-time hooks (use /setup-pre-commit).

helderberto/agent-skills · 70 tokens

performance-analysis

Comprehensive performance analysis, bottleneck detection, and optimization recommendations for Claude Flow swarms.

ruvnet/agentic-flow · 20 tokens

Verification & Quality Assurance

Comprehensive truth scoring, code quality verification, and automatic rollback system with 0.95 accuracy threshold for ensuring high-quality agent outputs and codebase reliability.

ruvnet/agentic-flow · 36 tokens

investigate-issues

Investigates identified issues in a LearningAgent session by reading the transcript, determining root causes, and updating issue files with investigation reports.

Unsupervisedcom/deepwork · 31 tokens

bugbash

Systematically explore and test any software project (CLI, API, Backend, Library, etc.) to find bugs, usability issues, and edge cases. Produces a structured report with full reproduction evidence (exact commands, inputs, logs, and tracebacks) for every issue.

av/facts · 58 tokens