Optimizing inference proxy for LLMs
OptiLLM is an OpenAI API-compatible proxy that applies inference-time techniques to improve the reasoning performance of large language models without training them. It is for developers and researchers who want to route existing model API calls through optimization methods.
Latest release v0.3.22 · 18 Jul 2026
These files are algorithmicsuperintelligence/optillm's own configuration. They tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
CLAUDE.md A 1,227 tok