ddalcu/mlx-serve

Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.

1.1kStars on the repository
3Mods indexed here, across every type
todayLast push, which is what freshness is scored on
noneNo LICENSE: all rights reserved, so bodies are not copied

bench

01

ddalcu/mlx-serve

Skill Claude CodeCodex

Skill "bench" from ddalcu/mlx-serve, covering benchmarking and comparison traps (these cost real days).

not rated 1.1k today A 53 tokens

release

02

ddalcu/mlx-serve

Skill Claude CodeCodex

Timings measured 2026-07-16 on the M4 Max 128 GB, AFTER the stopallengines port-wait fix (before it, everything below was 2.2× slower — see the gotcha in Benchmarking).

not rated 1.1k today A 42 tokens