Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/caiaffa/claude-code-ultimate-engineering-system/kubernetes-operabilitynpx skills add caiaffa/claude-code-ultimate-engineering-system --skill kubernetes-operabilitygit clone --depth 1 https://github.com/caiaffa/claude-code-ultimate-engineering-systemWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/caiaffa/claude-code-ultimate-engineering-system/kubernetes-operability)<a href="https://agentmods.dev/skills/caiaffa/claude-code-ultimate-engineering-system/kubernetes-operability"><img src="https://agentmods.dev/badge/skills/caiaffa/claude-code-ultimate-engineering-system/kubernetes-operability.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00026 | $0.00635 |
| Opus 5 | $0.00013 | $0.00318 |
| Sonnet 5 | $0.00005 | $0.00127 |
| Haiku 4.5 | $0.00003 | $0.00064 |
Grade A, and why
kubernetes-operability scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 62 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Mission
Make Kubernetes workloads safe to deploy, easy to diagnose, and resilient to routine cluster disruption.
When to use
- Reviewing K8s manifests or Helm charts.
- Diagnosing pod health or rollout issues.
- Validating autoscaling and runtime settings.
- Improving service operability.
Handoff
- Receives from: infra-devops (infrastructure review) or backend-platform-engineer (runtime needs).
- Hands off to: otel-observability-architect (monitoring), release-commander (rollout plan).
Probe rules
| Probe | Purpose | Common mistake |
|---|---|---|
readinessProbe |
"Can this pod serve traffic?" | Returns 200 before DB/cache connected |
livenessProbe |
"Is this pod stuck?" | Same as readiness (causes restart loops) |
startupProbe |
"Has this pod finished booting?" | Missing for slow-starting apps (liveness kills during boot) |
Rule: Readiness should check actual dependency availability. Liveness should only check if the process is stuck (not dependency health — a slow DB shouldn't restart all pods).
Resource settings
resources:
requests: # What the scheduler guarantees — base on p50 usage
cpu: 250m
memory: 256Mi
limits: # Hard ceiling — base on p99 + headroom
cpu: 1000m # Or omit CPU limit (throttling is worse than burst)
memory: 512Mi # Always set memory limit (OOM is better than node pressure)
Red flags
requests=limits(no burst room, constant throttling).- No
requestsset (scheduler can't make good decisions). - Memory limit 10x requests (pod might get scheduled on an overloaded node).
- HPA scaling on CPU when the bottleneck is I/O or queue depth.
- No PodDisruptionBudget on critical services.
terminationGracePeriodSecondsstill at default 30s for services that need longer shutdown.- Liveness probe with aggressive timeout that kills healthy-but-busy pods.
Graceful shutdown checklist
- SIGTERM received → stop accepting new requests.
- Finish in-flight requests (within
terminationGracePeriodSeconds). - Close database connections cleanly.
- Deregister from service discovery (readiness goes false).
- Exit 0.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 62 lines · 26 tokens per session scan A a85f1af49431
kubernetes-operability is a skill published in the GitHub repository caiaffa/claude-code-ultimate-engineering-system (17 stars, last pushed 2mo ago), licensed MIT. It adds 26 tokens to every session and 635 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
compute-env-setup
Set up a reproducible Feynman compute environment for research jobs. Use when a task needs Python/R packages, GPU libraries, containers, Modal, SSH, caches, or managed model runtime setup.
toolset-design
Design and validate new MCP toolsets, tools, and tool changes. Use when planning a new toolset, adding tools to an existing toolset, reviewing a toolset PR, or deciding whether a new tool is warranted. Covers the full lifecycle: eval-first validation, tool design methodology, consolidation patterns, and review…
nvca-chart-release
Release NVCA Operator chart changes from the native monorepo source to the vendored Helm chart. Use when updating the vendored NVCA Operator chart, changing NVCA image refs, publishing helm-nvca-operator, or validating the chart against a self-managed control plane.
nvca-values-customization
Customize NVCA Operator Helm chart values in the native monorepo. Use when modifying vendored defaults, changing stack-derived install values, adding deployment-time overrides, or updating scripts under deploy/helm/nvca-operator.
portainer-mcp-hygiene
How to drive the Portainer MCP server's tools correctly — both reading and mutating. Reading: project responses with select (JMESPath), where the heavy fields live (snapshots, status blocks, managed fields), how to handle non-JSON Docker/K8s proxy endpoints (container and pod logs, stats, exec), and how to interpret…
node-logs
Retrieve logs from a Kubernetes node — systemd units (journalctl) or files under /var/log. Use when you need node-level evidence: containerd, kubelet, kernel/OOM, or anything the pod's own log cannot show. Three access paths in order: hostscript (SSH), nodescript (debug pod), and localscript against the kubelet log…