agent observability skills

168 tagged agent observability, measured the same way as everything else here.

Browse within: ci-cd 43clickhouse 43evals 43ai-monitoring 32agent-monitoring 31ai-error-monitor 30ai-observability 30llm-observability 30agent-builder 25agent-orchestration 25agent-workspace 25data-observability 20Apple Silicon 12agent-inventory 12

agent-release-gate

01

Agenta-AI/agenta

Skill Claude CodeCodex

Run the agent release gate — a portable, wire-level QA harness for the agent runtime. Drives the same product endpoint the playground drives and asserts on the SSE frame stream and real side effects, never on model prose, so it works against any deployment (cloud or self-hosted) from three env vars. Use before an…

4.6k 2d ago A 126 tokens

Agenta-AI/agenta

Skill Claude CodeCodex

Where to put frontend code (package vs app layer) and how to use the @agenta/ packages. Use when authoring or moving code in web/packages, choosing between @agenta/ui, @agenta/entities, @agenta/entity-ui, @agenta/shared, @agenta/playground, using molecules, loadable/runnable bridges, the EntityPicker, or writing…

4.6k 2d ago A 88 tokens

Agenta-AI/agenta

Skill Claude CodeCodex

Use this skill to create and publish changelog announcements for new features, improvements, or bug fixes. This skill handles the complete workflow - creating detailed changelog documentation pages, adding sidebar announcement cards, and ensuring everything follows project standards. Use when the user mentions adding…

4.6k 2d ago A 75 tokens

api-endpoints

04

latitude-dev/latitude-llm

Skill Claude CodeCodex

Adding or changing API operations in @repo/operations. One source of truth (defineOperation + a Zod schema) becomes an HTTP endpoint, an OpenAPI operation, an MCP tool, TS + Python SDK methods, a latitude CLI command, and an in-process agent tool — descriptions and contracts must be written with all of these readers…

4.6k 3d ago A 78 tokens original MIT

fix-datadog-issues

05

latitude-dev/latitude-llm

Skill Claude CodeCodex

Find, triage, and fix production errors captured by Datadog Error Tracking, then open a PR. Use when asked to look at "Datadog issues/incidents/errors", "find and fix bugs from Datadog", investigate the most-frequent or newest production errors, or work a specific Datadog Error Tracking issue.

4.6k 3d ago A 76 tokens original MIT

web-frontend

06

latitude-dev/latitude-llm

Skill Claude CodeCodex

When to use: apps/web UI — routes, @repo/ui, TanStack Start server functions and collections, navigation (Link vs useNavigate), forms (useForm with createFormSubmitHandler + fieldErrorsAsStrings when Zod validation errors should appear on fields), Tailwind layout rules, design-system updates, and useEffect /…

4.6k 3d ago A 67 tokens original MIT

mintlify

07

FailproofAI/failproofai

Skill Claude CodeCodex

Build and maintain documentation sites with Mintlify. Use when creating docs pages, configuring navigation, adding components, or setting up API references.

1.7k +104 today A 30 tokens

fp-cloud-cli

08

FailproofAI/failproofai

Skill Claude CodeCodex

The way to answer "how are my production AI agents doing?" and to run the team's agent-observability deployment — reach for it even on casual phrasing that names no tool. Trigger when the user wants to: • inspect agent telemetry — did agents error/fail/go flaky; sessions, events, latency, token usage, slowest models…

1.7k +104 today A 245 tokens

compose-graphics

09

Jwuthri/Tracely-ai

Skill Claude CodeCodex

Advanced Compose visuals - Material 3 Expressive motion physics, AGSL shaders (Android 13+), Canvas/DrawScope generative, graphicsLayer effects.

1.2k +2 3d ago A 36 tokens original MIT

Jwuthri/Tracely-ai

Skill Claude CodeCodex

Compose Multiplatform / KMP patterns - expect/actual composables, platform-specific code, density and font handling cross-target, iOS/Android/Desktop interop.

1.2k +2 3d ago A 40 tokens original MIT

paint

11

Jwuthri/Tracely-ai

Skill Claude CodeCodex

Paint a complete visual universe with genjutsu - art direction brainstorm, design system, implementation, audit. Anti-AI-slop design pipeline. Adapts to Web, Android (Compose), Apple (SwiftUI).

1.2k +2 3d ago A 46 tokens original MIT

agentacct-workflow

12

mikehasa/agentacct

Skill Claude CodeCodex

Use when working in a repo with agentacct MCP configured, or when asked to track coding-agent work, smoke-test agentacct integrations, or report objective AI-agent task evidence.

675 3d ago A 40 tokens original MIT

code-review

13

sandbaseai/sandbase-harness

Skill Claude CodeCodex

Reviews a supplied code path or diff for correctness, security, maintainability, and style without executing or modifying it.

638 2d ago A 25 tokens original Apache-2.0

reins

14

pegasi-ai/reins

Skill Claude CodeCodex

Use whenever security, policies, governance, guardrails, compliance, or safety are relevant — including blocked commands, audit trails, dangerous operations, deletions, file modifications, shell commands, MCP access, API calls, network requests, credentials, or any action that could be irreversible or destructive.

408 3mo ago C 58 tokens original Apache-2.0

reins

15

pegasi-ai/reins

Skill Claude CodeCodex

Use this skill whenever security, policies, governance, guardrails, compliance, or safety are relevant — including blocked commands, audit trails, dangerous operations, deletions, file modifications, shell commands, MCP access, API calls, network requests, credentials, or any action that could be irreversible or…

408 3mo ago C 90 tokens original Apache-2.0

bitrouter

16

bitrouter/bitrouter

Skill Claude CodeCodex

Use this skill when the user wants to install, configure, run, or troubleshoot BitRouter — an LLM proxy that runs two ways: a local Rust daemon at http://localhost:4356 (BYOK) or BitRouter Cloud at https://api.bitrouter.ai/v1 (managed, brk keys, Stripe credits or x402 wallet). Unifies OpenAI, Anthropic, Google…

222 3d ago C 261 tokens original Apache-2.0

bitrouter/bitrouter

Skill Claude CodeCodex

Use when evaluating BitRouter route decisions or Eval Exchange subjects with task-native verifiers, human reviewers, private enterprise evaluators, agentic judges, or genuinely uncategorized evaluator sources.

222 3d ago A 44 tokens original Apache-2.0

bitrouter/bitrouter

Skill Claude CodeCodex

Use when a user wants to run, compare, resume, audit, share, or submit a Harbor benchmark through BitRouter, including choosing a Harbor dataset and agent, confirming routed providers and models, or optionally operating BitRouter OSS on AWS.

222 3d ago A 54 tokens original Apache-2.0

ax-repo

19

Necmttn/ax

Skill Claude CodeCodex

Star the ax repo, file an issue / bug report, or fork-and-open-a-PR against github.com/Necmttn/ax on the user's behalf, by shelling out to the gh CLI. Triggers when the user says "star ax", "star the repo", "I want to support ax", "report this as an ax bug", "file an ax issue", "open an issue on ax", "this looks like…

104 6d ago A 211 tokens AGPL-3.0

dojo

20

Necmttn/ax

Skill Claude CodeCodex

Surplus-quota training loop over the ax graph - the agent burns the remaining 5h/7d plan-quota window on self-improvement: locking pending verdicts, filling briefs, backtesting routing classes, minting proposals, running worktree experiments, and drafting upstream issue reports. Triggers when the user says "/dojo"…

104 6d ago A 0 tokens AGPL-3.0

efficient-dispatch

21

Necmttn/ax

Skill Claude CodeCodex

Model-routing orchestration for any expensive frontier model (Fable, Opus, GPT-5.x) - the main model keeps judgment and Q&A review, mechanical subagent dispatches carry an explicit cheaper model, and ax measures whether the routing actually worked. Use when orchestrating codebase-heavy or token-heavy work with…

104 6d ago A 148 tokens AGPL-3.0

monte-carlo-data/mc-agent-toolkit

Skill Claude CodeCodex

Investigate data incidents and find root causes using Monte Carlo's observability data. Guides the agent through systematic investigation: alert lookup, lineage tracing, ETL checks, query analysis, and data profiling. Activates when a user asks about data issues, incidents, alerts, or why data looks wrong.

91 8d ago A 73 tokens original Apache-2.0