intentee/paddler

Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale ๐Ÿ“๐Ÿฆ™ Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.

About the project

Paddler is an open-source service that runs and distributes large language model and vision-language model inference across self-managed CPU or GPU infrastructure. It is for product, DevOps, and LLM operations teams that need private, scalable model serving, embeddings, and more predictable costs. The catalogue add-ons help agents operate Paddler.

Latest release v4.1.0 โ€” v4.1.0 (OpenCode support) ยท 19 Jul 2026

These files are intentee/paddler's own configuration. They tell Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

1,667Stars on the repository
3Files it configures its agents with
103Tokens loaded in every session
1Agent configured

Instructions

Skills