Jacki1223/inference-autopilot

Evidence-driven SGLang inference deployment optimization.

141Stars on the repository
1Mods indexed here, across every type
7d agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

inference-autopilot

01

Jacki1223/inference-autopilot

Skill Claude CodeCodex

Analyze, benchmark, diagnose, and optimize large-model inference deployments from hardware inventory, model details, workload traces, and latency or throughput SLOs. Use when Codex needs to tune SGLang launch parameters, run bounded single-host GPU experiments, plan deployment topology, inspect GPU or CPU profiles…

not rated 141 7d ago A 98 tokens original Apache-2.0