MCP server Claude CodeCodexCursor
Governed AI-ops for GPU inference clusters (vLLM + Ray Serve/Jobs): latency/utilization RCA, replica scaling, drain, model lifecycle, and destructive-op guardrails with a built-in governance harness (audit, budget, undo, risk tiers). Runs locally from the inference-aiops Python package.
0 6d ago A
tokens not measured
original MIT