JoniMartin27/inferbench

Benchmark local LLM inference speed (tokens/sec) on your own hardware — llama.cpp native + cloud APIs, 124-model catalog, optimal-quant picker, and an MCP serve mode.

2Stars on the repository
2Mods indexed here, across every type
7d agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

JoniMartin27/inferbench

Instructions file

Claude Code instructions for JoniMartin27/inferbench, covering instrucciones para claude code en este proyecto, antes de tocar nada, lo que no debes hacer (load-bearing), schemas de optimización son por motor, no uniformes and no simules motores.

2 7d ago A 4,281 tokens original MIT