observability

observability is a skill for Claude Code, Codex from aethrox/doctrine. It costs 100 tokens per session (1,120 once invoked), scanned A, original, MIT.

A set of practices for monitoring a service through its logs, measurements, and request traces. It defines health using latency, errors, traffic, capacity use, and agreed targets.

In plain words
What is it for?
Adding structured logs, metrics, and tracing to a service. Defining service targets, error budgets, and alerts that link to runbooks.
Why use it?
It lets people tell whether a service is working without reading its code. It also helps connect failures to the relevant requests and operating instructions.

Skill for Claude CodeCodex

Part of the doctrine plugin — 33 skills shipped together

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/aethrox/doctrine/observability
Any agent
npx skills add aethrox/doctrine --skill observability
Clone the repo
git clone --depth 1 https://github.com/aethrox/doctrine

Made for: Claude Code, Codex.

Or install doctrine, the plugin that ships this one along with the rest of its 33 skills.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for observability

README.md
[![agentmods](https://agentmods.dev/badge/skills/aethrox/doctrine/observability.svg)](https://agentmods.dev/skills/aethrox/doctrine/observability)
Your own site
<a href="https://agentmods.dev/skills/aethrox/doctrine/observability"><img src="https://agentmods.dev/badge/skills/aethrox/doctrine/observability.svg" alt="Measured on agentmods" height="20"></a>
Per session 100 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 1,120 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00100 $0.01120
Opus 5 $0.00050 $0.00560
Sonnet 5 $0.00020 $0.00224
Haiku 4.5 $0.00010 $0.00112

Measured 4d ago against content hash 21e8e935a757, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

observability scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

skills/observability/SKILL.md · 51 lines

How it starts

The opening of the file, as written. The whole thing — 51 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Observability

A service is not done until someone who has never read its code can answer "is it healthy right now" from its telemetry alone. That means structured logs that a platform can index and correlate, the four golden signals tracked per service, and an explicit SLO with an error budget, not "we'll notice if it breaks."

Phase 1: Structured logging

  • Every log line is a structured record (JSON or the platform's structured format), not a hand-built string; fields are queryable and indexable; a concatenated string is only greppable.
  • Every log line inside a request/job carries a correlation ID (trace ID / request ID) so a slow request, its trace, and its log lines can be joined without copy-pasting timestamps.
  • Log levels mean something consistent across the codebase: error = needs human attention, warn = degraded but self-recovering, info = significant state change (request completed, job started), debug = development-time detail, off by default in production.
  • Never log a secret, full payload, or PII field directly; log its presence or a redacted/hashed form (same redaction discipline as diagnosing-bugs and secure-coding).
  • Log the outcome, not just the attempt: "started X" without a matching "X succeeded / X failed (reason)" leaves a gap no dashboard can fill in later.

Phase 2: The four golden signals

Every service exposes these four, regardless of what else it tracks:

Signal What it measures Typical source
Latency time to serve a request: track success and failure latency separately, a fast failure skews the average down and hides a real slowdown request duration histogram
Errors rate of failed requests (5xx, exceptions, failed jobs) error counter / request outcome
Traffic demand on the system (requests/sec, queue depth, jobs/min) request or throughput counter
Saturation how full the constrained resource is (CPU, memory, connection pool, queue length) resource utilization gauge

Read the full file on GitHub · 51 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 4d ago First seen · 51 lines · 100 tokens per session scan A 21e8e935a757

Subscribe to this mod's changes

observability is a skill published in the GitHub repository aethrox/doctrine (18 stars, last pushed 22d ago), licensed MIT. It adds 100 tokens to every session and 1,120 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

debug-optimize-lcp

Guides debugging and optimizing Largest Contentful Paint (LCP) using Chrome DevTools MCP tools. Use this skill whenever the user asks about LCP performance, slow page loads, Core Web Vitals optimization, or wants to understand why their page's main content takes too long to appear. Also use when the user mentions…

ChromeDevTools/chrome-devtools-mcp · 99 tokens

specflow-use

To connect Rosetta with Grid Dynamics SpecFlow MCP; only when SpecFlow is mentioned and the MCP is installed.

griddynamics/rosetta · 27 tokens

opik-diagnose

Surface the Opik traces worth a developer's attention, ranked by signal — errors, failed tool calls, latency, regressions, and low online-eval scores — plus Diagnostics issues. Reads live/production traces via the SDK (searchtraces and agentinsights) and works with no MCP; uses the MCP issue entity when connected.…

comet-ml/opik-mcp · 147 tokens

ue-mcp-epic-routing

Use when deciding between ue-mcp's native category actions and Epic's wrapped ToolsetRegistry tools (the epic actions, incl. the Blueprint graph DSL) for a task in Unreal. Pulls in when authoring Blueprint graph bodies, or any time both a native action and an epic action could do the job and you need to pick.

db-lyon/ue-mcp · 81 tokens

mcp-google-map-project

Project knowledge for developing and maintaining @cablate/mcp-google-map. Architecture, Google Maps API guide, GIS domain knowledge, and design decisions. Read this skill to onboard onto the project or make informed development decisions.

cablate/mcp-google-map · 49 tokens

patsnap-maintain-payment

Patsnap Maintain Payment MCP for AI agents. A global patent search tool for freedom-to-operate assessment. It supports query-based, semantic, and image search, and can be combined with patent family, legal status, and analytics views to help users identify patents that may create FTO barriers more systematically. It…

patsnap/mcp · 84 tokens