ops-debugging

A placeholder runbook for diagnosing and operating the datatorag-mcp production gateway, including plugin discovery, OAuth token failures, containers, and health checks. A runbook is a set of operational instructions.

In plain words
What is it for?
Use it when updating or rediscovering plugins, investigating OAuth failures, checking containers and service health, or following the gateway's production troubleshooting process.
Why use it?
It points to the procedures and private references needed to investigate production problems without copying sensitive infrastructure details into a public file.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/datatorag/mcp-gateway/ops-debugging
Any agent
npx skills add datatorag/mcp-gateway --skill ops-debugging
Clone the repo
git clone --depth 1 https://github.com/datatorag/mcp-gateway

Made for: Claude Code, Codex.

Per session 50 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 2,972 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00050 $0.02972
Opus 5 $0.00025 $0.01486
Sonnet 5 $0.00010 $0.00594
Haiku 4.5 $0.00005 $0.00297

Measured 2d ago against content hash b859cd860e48, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

ops-debugging scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

step, rebuild command, health-check curl, plugin reinstall exec, log-tailing,
.claude/skills/ops-debugging/SKILL.md · 220 lines

How it starts

The opening of the file, as written. The whole thing — 220 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Ops Debugging — Production Gateway

Diagnose and operate the production gateway without hard-coding any infrastructure details in this file. This skill lives in a public repo; live values (host, SSH key, ports beyond what's already public in the compose files, profile names) come from private memory or the deploy/db-query skills, never from here.

Sources of truth

Don't duplicate these — read them first, then come back here for the gap:

  • Deploy skill (.claude/skills/deploy/SKILL.md) — SSH access, .env render step, rebuild command, health-check curl, plugin reinstall exec, log-tailing, its own 5-item Troubleshooting section.
  • db-query skill (.claude/skills/db-query/SKILL.md) — how to query prod (Neon, via Neon MCP) vs local dev (docker exec psql), safety rails, canned recipes including the tool-registry query.
  • Memory refs for live values: reference_mcp_gateway_instance (host/region), reference_plugin_registry (installed plugins, reinstall notes), reference_neon_database (prod project id/region).

This skill only adds what those don't cover: the full plugin re-discovery recipe (deploy skill stops at "see reference_plugin_registry for details"), a symptom-first failure-mode table, and verification patterns that combine health checks + DB state.

Plugin update + tool re-discovery

When a plugin's tool set changes (new/renamed/removed tools) and a rebuild alone won't fix the tools table, run the full re-discovery recipe. Proven in prod 2026-07-18.

  1. Pull + build the plugin inside its running container (see deploy skill step 5 for the git pull && pnpm install && npx tsc exec).

  2. Restart the gateway so the plugin child process picks up the new build.

  3. Run a re-discovery script from inside the gateway container — it must execute from /app/apps/gateway (module resolution for @modelcontextprotocol/sdk and the postgres driver fails from /tmp or other paths):

    docker exec -i <gateway-container> bash -c \
      'cat > /app/apps/gateway/rediscover.mjs && cd /app/apps/gateway && node rediscover.mjs; rm -f /app/apps/gateway/rediscover.mjs' < rediscover.mjs
    

Read the full file on GitHub · 220 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 2d ago First seen · 220 lines · 50 tokens per session scan A b859cd860e48

Subscribe to this mod's changes

ops-debugging is a skill published in the GitHub repository datatorag/mcp-gateway (3 stars, last pushed 2d ago), licensed MIT. It adds 50 tokens to every session and 2,972 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

deploy-docker-compose

Run the Omnigent server as a Docker compose stack (server + Postgres) on any Docker host — your laptop, a VPS, EC2 by hand, or as the base layer of any container-platform deploy. Invoke when the user wants to build the image, bring up the compose stack, debug the stack on a host they already have, or extend the stack…

omnigent-ai/omnigent · 84 tokens

reproduce-issue

The single skill for reproducing an nx issue. Given a GitHub issue number (human entry) OR explicit repro parameters (agent entry), it runs the reproduction ENTIRELY inside an isolated Docker sandbox — gVisor on Linux, the Docker VM on macOS — so the untrusted repro's install scripts and commands never execute on the…

nrwl/nx · 91 tokens

timoni

Use when deploying applications to Kubernetes with Timoni. Covers installing and upgrading module instances from OCI registries, composing multi-app deployments with bundles, injecting values from clusters or CI with runtimes, targeting multiple clusters, and authoring, testing, signing and publishing modules with CUE.

stefanprodan/timoni · 59 tokens

atmos-aws-ecr

AWS ECR commands in Atmos: atmos aws ecr login, ECR auth integrations, Docker credential writes, registry login via identity or explicit registry.

cloudposse/atmos · 36 tokens

windows-builder

Build Windows images with Packer using WinRM communicator and PowerShell provisioners. Use when creating Windows AMIs, Azure images, or VMware templates.

hashicorp/agent-skills · 33 tokens

release-openclaw-plugin-testing

Plan and run pre-release OpenClaw plugin validation across bundled plugins, package artifacts, lifecycle commands, doctor/fix, config round-trip, gateway startup, SDK compatibility, Docker E2E, Package Acceptance, and Testbox proof.

openclaw/openclaw · 54 tokens