bench
01Skill Claude CodeCodex
Skill "bench" from ddalcu/mlx-serve, covering benchmarking and comparison traps (these cost real days).
not rated 1.1k today A 53 tokens
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
Skill Claude CodeCodex
Skill "bench" from ddalcu/mlx-serve, covering benchmarking and comparison traps (these cost real days).
Skill Claude CodeCodex
Timings measured 2026-07-16 on the M4 Max 128 GB, AFTER the stopallengines port-wait fix (before it, everything below was 2.2× slower — see the gotcha in Benchmarking).