JoniMartin27/inferbench

Benchmark local LLM inference speed (tokens/sec) on your own hardware — llama.cpp native + cloud APIs, 124-model catalog, optimal-quant picker, and an MCP serve mode.

2Stars on the repository
2Mods indexed here, across every type
9d agoLast push, which is what freshness is scored on
MITLicence, which decides whether bodies are shown

JoniMartin27/inferbench

Skill Claude CodeCodex

Build, launch, smoke-test and drive the InferBench FastAPI backend (:7777) and its inference-engine orchestration. Use when asked to run, start, launch, boot, smoke-test, benchmark, or drive the inferbench backend / API / engines (llama.cpp, ollama, vLLM, SGLang, TGI), or to verify an engine actually starts and runs a…

not rated 2 9d ago A 90 tokens original MIT