Instructions file
Instructions for intentee/paddler: When working with this codebase, prioritize readability over cleverness. Ask clarifying questions before making architectural changes.
1.7k 1mo ago A 103 tokens
original Apache-2.0
Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.
Instructions file
Instructions for intentee/paddler: When working with this codebase, prioritize readability over cleverness. Ask clarifying questions before making architectural changes.