Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add gke-labs/kube-agents --skill gke-tpu-dynamic-slices-monitoringgit clone --depth 1 https://github.com/gke-labs/kube-agentsWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/gke-labs/kube-agents/gke-tpu-dynamic-slices-monitoring)<a href="https://agentmods.dev/skills/gke-labs/kube-agents/gke-tpu-dynamic-slices-monitoring"><img src="https://agentmods.dev/badge/skills/gke-labs/kube-agents/gke-tpu-dynamic-slices-monitoring/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/gke-labs/kube-agents/gke-tpu-dynamic-slices-monitoring"><img src="https://agentmods.dev/badge/skills/gke-labs/kube-agents/gke-tpu-dynamic-slices-monitoring.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00103 | $0.02042 |
| Opus 5 | $0.00051 | $0.01021 |
| Sonnet 5 | $0.00021 | $0.00408 |
| Haiku 4.5 | $0.00010 | $0.00204 |
Grade A, and why
gke-tpu-dynamic-slices-monitoring scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 9d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
This is a copy
86% identical to gke-ai-troubleshooting-tpu-dynamic-slices-monitoring — 53 lines differ, which has more behind it and is treated as the original. This page carries a canonical link to it rather than competing with it.
How it starts
The opening of the file, as written. The whole thing — 189 lines — stays where its author put it; the contents beside it link to each section on GitHub.
GKE TPU Dynamic Slices Monitoring & Management
Monitors the status of TPU Slice custom resources, troubleshoots provisioning failures, validates workload manifests on dynamic slices, and performs cleanups.
Prerequisites
- Cloud Logging enabled for the project.
kubectlandgcloudCLIs configured to access the GKE cluster.
Diagnostic Workflow
Step 0: Context Acquisition & Time Window Definition
Gather project, cluster, and slice context using cluster tools or the following parameters:
- Project ID:
{project_id}(e.g.,my-gcp-project) - Cluster Name:
{cluster_name}(e.g.,tpu-cluster) - Region/Zone:
{location}(e.g.,us-central1-a) - Slice Name:
{slice_name}(e.g.,test-slice) - Issue Time:
{timestamp}(Optional; default to the last 30 minutes window[T - 30m]to[T + 30m])
Step 1: Describe the Slice Custom Resource [Low Risk]
When asked to inspect, troubleshoot, or check a slice status, immediately execute kubectl describe slice {slice_name} using available cluster tools to perform the inspection. Parse the resulting Status.Conditions output against the condition table below to diagnose the exact state and provide concrete recommendations.
-
Command:
kubectl describe slice {slice_name}
State & Reason Analysis
Analyze the Status.Conditions (especially Type: Ready and its Reason and
Status):
| Lifecycle State / Reason | Meaning | Recommended Action |
|---|---|---|
SliceNotCreated |
GKE Slice Controller | Wait a few minutes and |
| : : is initializing the : re-check slice status. : | ||
| : : slice and performing : : | ||
| : : resource checks. : : | ||
SliceCreationFailed |
Prerequisites | Verify selected nodes |
| : : validation failed : exist, are unallocated, : | ||
| : : (e.g., selected nodes : and topology matches : | ||
| : : don't exist, nodes are : partition count. : | ||
| : : already used by : : | ||
| : : another slice, or the : : | ||
| : : topology doesn't match : : | ||
| : : the number of : : | ||
| : : partitions). : : | ||
ACTIVATING |
GKE is actively | Monitor node |
| : : forming and : provisioning. : | ||
| : : provisioning the TPU : : | ||
| : : slice. : : | ||
ACTIVE |
The TPU slice is | Proceed to deploy or |
| : : successfully formed : check workloads. : | ||
| : : and ready to host : : | ||
| : : workloads. : : | ||
ACTIVE_DEGRADED |
The slice is usable, | Monitor workload logs |
| : : but one or more : for interconnect or : | ||
| : : sub-blocks are : device errors. Check : | ||
| : : degraded. : faulty node VMs. : | ||
FAILED |
GKE failed to form the | Ensure all selected |
| : : TPU slice (e.g., : nodes belong to the : | ||
| : : selected nodes are not : same reservation block. : | ||
| : : part of the same : : | ||
| : : reservation block). : : | ||
DEACTIVATING |
The slice is | Wait for dismantling to |
| : : dismantling (triggered : finish, or patch : | ||
| : : by user deletion or a : finalizers if stuck. : | ||
| : : critical systemic : : | ||
| : : failure). : : | ||
INCOMPLETE |
The terminal phase | No action required; the |
| : : before the Slice CR is : resource will be : | ||
| : : deleted from the : removed shortly. : | ||
| : : cluster. : : |
What ships with it
2 files beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 9d ago First seen · 189 lines · 103 tokens per session scan A f70e877fff66
gke-tpu-dynamic-slices-monitoring is a skill published in the GitHub repository gke-labs/kube-agents (54 stars, last pushed today), licensed Apache-2.0. It adds 103 tokens to every session and 2,042 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. It is 86% identical to gke-ai-troubleshooting-tpu-dynamic-slices-monitoring, differing in 53 lines, and is treated as a copy.
Other skills, from other repositories
gke-compute-classes
Configures, optimizes, and troubleshoots GKE ComputeClasses. Use when configuring Spot VMs with on-demand fallback, targeting specific accelerators (GPUs/TPUs) or machine families, restricting ComputeClass access, or debugging pending pods related to node pool auto-creation. Do not use for cluster-level Node Auto…
jetson-diagnostic
Read-only Jetson health snapshot for identity, memory, GPU, thermal, power, storage, services, and top processes.
doca-socket-relay
Use this skill when the operator is driving the DOCA Socket Relay to bridge a socket-oriented host application onto a BlueField DPU peer without rewriting it — picking the deployment shape (in-process, sidecar, or BlueField service container), configuring the host-side socket and the DPU-side forwarding endpoint…
offensive-z-wave
Z-Wave attack methodology — sniffing with Z-Force / EZ-Wave / RTL-SDR + ZniffMobile, S0 (legacy) network-key derivation flaw and key reuse, S2 (modern) ECDH commissioning analysis, replay/injection on unauthenticated nodes, default-key brute-force on test deployments, and home-automation hub pivots. Use when targeting…
hsb-flash
Flash the FPGA on an HSB board connected to an NVIDIA devkit. Supports HSB Lattice boards (FPGA versions 2407, 2412, 2507, 2510) and Leopard Imaging VB1940 "all-in-one" cameras (FPGA versions 2507, 2510). Uses release-specific YAML manifests and board-type-specific program commands. Lattice and VB1940 commands must…
jetson-validate-image
Use after jetson-flash-image to run static BSP checks, on-target smoke/regression tests on a flashed DUT, or both. Not for build or flash steps. Triggers: validate bsp, on-target validation.