Benchmark local LLM inference speed (tokens/sec) on your own hardware — llama.cpp native + cloud APIs, 124-model catalog, optimal-quant picker, and an MCP serve mode.
Latest release v0.3.0 · 19 Aug 2026
These files are JoniMartin27/inferbench's own configuration. They tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
CLAUDE.md A 4,281 tok