vLLM-backed self-hosted inference server — OpenAI-compatible API, model registry, request observability, SSE streaming, and a Textual TUI playground for real-time model comparison.
Latest release v0.6.0 — v0.6.0 — Phase B: Admission Bounding, Batch Queueing, Multi-Model Process Split · 10 Aug 2026
2 files for Codex, OpenCode and Claude Code: inference-x AGENTS.md, inference-x CLAUDE.md — 3,257 tokens loaded in every session.
AGENTS.md A 2,642 tok CLAUDE.md A 615 tok These files are coeusyk/inference-x's own configuration — they tell Codex, OpenCode and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.