coeusyk/inference-x

vLLM-backed self-hosted inference server — OpenAI-compatible API, model registry, request observability, SSE streaming, and a Textual TUI playground for real-time model comparison.

Latest release v0.6.0 — v0.6.0 — Phase B: Admission Bounding, Batch Queueing, Multi-Model Process Split · 10 Aug 2026

2 files for Codex, OpenCode and Claude Code: inference-x AGENTS.md, inference-x CLAUDE.md — 3,257 tokens loaded in every session.

2Stars on the repository
2Files it configures its agents with
3,257Tokens loaded in every session
3Agents configured

Instructions

These files are coeusyk/inference-x's own configuration — they tell Codex, OpenCode and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.