k8s-incident-triage

k8s-incident-triage is a skill for Claude Code, Codex from initializ/forge. It costs 41 tokens per session (1,510 once invoked), scanned A, original, Apache-2.0.

A read-only troubleshooting workflow for Kubernetes, a system for running containerised applications. It examines workloads, pods, namespaces, rollout status, and logs, then gives possible causes, supporting evidence, and next commands.

In plain words
What is it for?
Use it to investigate a namespace, deployment, pod, workload, selector, or rollout using plain-language or structured requests.
Why use it?
It helps investigate failed deployments, pending pods, crash loops, and unhealthy services without changing the cluster. The evidence and hypotheses make the problem easier to narrow down.

Skill for Claude CodeCodex

Written for no agent in particular: nothing here depends on one.

Good fit Use it to investigate a namespace, deployment, pod, workload, selector, or rollout using plain-language or structured requests.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/initializ/forge/k8s-incident-triage
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add initializ/forge --skill k8s-incident-triage
Clone the repo
git clone --depth 1 https://github.com/initializ/forge

Made for: Claude Code, Codex.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for k8s-incident-triage

README.md
[![agentmods](https://agentmods.dev/badge/skills/initializ/forge/k8s-incident-triage.svg)](https://agentmods.dev/skills/initializ/forge/k8s-incident-triage)
Your own site
<a href="https://agentmods.dev/skills/initializ/forge/k8s-incident-triage"><img src="https://agentmods.dev/badge/skills/initializ/forge/k8s-incident-triage.svg" alt="Measured on agentmods" height="20"></a>
Per session 41 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,510 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. A grade says what 26 rules found in the file — not that it is safe. Third-party audits
  • NVIDIA SkillSpector warn 7 Sept 2026
SkillSpector: 1 finding, up to high

These are SkillSpector’s own severities. On a checked sample its high-severity flags on skills were ~96% false positives — a documented command, a public API, a “never do X” rule — so we show them as a caution to read, not a verdict. Why →

  • high Privilege Escalation · line 24
    Code accesses credential files (SSH keys, AWS credentials, etc.). This could indicate credential theft attempts.
    Fix: Remove references to credential paths. Use environment variables or secrets managers. For docs, use placeholder paths (e.g., /path/to/config). Never load .env or token files in production code paths.
How audits are shown
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00041 $0.01510
Opus 5 $0.00020 $0.00755
Sonnet 5 $0.00008 $0.00302
Haiku 4.5 $0.00004 $0.00151

Measured 8d ago against content hash bcd28d2708bc, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-08, from the pricing page.

Security

Grade A, and why

k8s-incident-triage scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 8d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

forge-skills/local/embedded/k8s-incident-triage/SKILL.md · 322 lines

How it starts

The opening of the file, as written. The whole thing — 322 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Kubernetes Incident Triage

Performs read-only triage of Kubernetes workloads and namespaces using kubectl.

Supports:

  • Natural language input (human mode)
  • Structured JSON input (automation mode)

This skill NEVER mutates cluster state.


Tool Usage

This skill uses cli_execute with kubectl commands exclusively. NEVER use http_request or web_search to interact with Kubernetes. All cluster operations MUST go through kubectl via the cli_execute tool.


Tool: k8s_triage

Diagnose unhealthy Kubernetes workloads, pods, or namespaces.

Output format: Use markdown tables for pod/workload status summaries. Wrap kubectl output and log excerpts in text code blocks. Use bash for recommended next-step commands.


Input Modes

1) Human Mode (Natural Language)

Input is a plain string.

Examples:

  • triage payments-prod
  • triage deployment payments-api in payments-prod
  • why are pods pending in checkout-prod?
  • investigate crashloop in payments-prod
  • triage pod api-7c9f6d7f86-abcde in payments-prod
  • check rollout of deployment payments-api in prod

Behavior:

  • Parse namespace, workload, pod, or selector intent.
  • If namespace omitted, use $DEFAULT_NAMESPACE if set.
  • If ambiguity exists, default to namespace-level triage.
  • Never require the user to remember JSON fields.

2) Automation Mode (Structured JSON)

Input JSON schema:

{ "namespace": "payments-prod", "workload_kind": "deployment", "workload_name": "payments-api", "pod_name": null, "label_selector": null, "include_logs": true, "logs_tail_lines": 200, "include_previous_logs": true, "events_limit": 50, "include_node_diagnostics": true, "include_metrics": false, "output_format": "markdown" }

Rules:

  • namespace is required.
  • If pod_name provided → pod-level triage.
  • If workload fields provided → workload-level triage.
  • Else → namespace scan.

Triage Process

Step 0 — Preconditions

Verify cluster access:

kubectl version --client kubectl cluster-info

Read the full file on GitHub · 322 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 8d ago First seen · 322 lines · 41 tokens per session scan A bcd28d2708bc

Subscribe to this mod's changes

k8s-incident-triage is a skill published in the GitHub repository initializ/forge (227 stars, last pushed yesterday), licensed Apache-2.0. It adds 41 tokens to every session and 1,510 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.