Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale ๐๐ฆ Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.
Paddler is an open-source service that runs and distributes large language model and vision-language model inference across self-managed CPU or GPU infrastructure. It is for product, DevOps, and LLM operations teams that need private, scalable model serving, embeddings, and more predictable costs. The catalogue add-ons help agents operate Paddler.
Latest release v4.1.0 โ v4.1.0 (OpenCode support) ยท 19 Jul 2026
These files are intentee/paddler's own configuration. They tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.
CLAUDE.md A 103 tok .claude/skills/running-all-tests/SKILL.md A 47 tok .claude/skills/running-coverage/SKILL.md A 42 tok