FlashML-org/FreeToken

FreeToken brings datacenter-scale model serving to your desktop. Run massive models locally, fast and efficiently.

About the project

FreeToken is an inference-serving engine, meaning software that runs machine-learning models and responds to requests, for running large open-weight mixture-of-experts models on personal hardware. It coordinates GPUs, CPUs, system memory, and their connections so developers can use these models locally through APIs compatible with Anthropic and OpenAI clients.

Latest release v0.1.2 · 19 Aug 2026

These files are FlashML-org/FreeToken's own configuration. They tell Codex, OpenCode and Claude Code how to work on this repository, so they are not mods to install elsewhere. Copy one as a starting point and replace the parts that are about this project.

11,783Stars on the repository
2Files it configures its agents with
988Tokens loaded in every session
3Agents configured

Instructions