inference agents

18 tagged inference, measured the same way as everything else here.

Browse within: compilers 9gpu-computing 9llm-inference 9llmops 9macos 9Apple Silicon 8fastapi 8local-llm 8

claude-code

01

raullenchai/Rapid-MLX

Agent

Point Anthropic's Claude Code at a local rapid-mlx server. Claude Code speaks the Anthropic Messages API (POST /v1/messages); rapid-mlx implements that route natively, so you can drive Claude Code with any local model.

3.6k 2d ago A 0 tokens

matrix

02

raullenchai/Rapid-MLX

Agent

This page renders the Tier-1 agent × model-family integration matrix truthfully from the authoritative test suite in tests/integrations/ — the matrix cells (testagentsmatrix.py), the family aliases and strict-xfail rules (conftest.py), and the pilot run recorded in tests/integrations/README.md.

3.6k 2d ago A 0 tokens

openhands

03

raullenchai/Rapid-MLX

Agent

Point OpenHands (formerly OpenDevin) at a local rapid-mlx server. OpenHands drives its CodeActAgent inside a Docker sandbox and reaches the model over the OpenAI-compatible chat completions API (POST /v1/chat/completions) via LiteLLM.

3.6k 2d ago A 0 tokens

task-agents

04

matthiasn/lotti

Agent

The primary agent workflow — inference setup resolution, the automation switch, evidence-first execution, tool policy, and the proposal/confirmation loop.

1.2k 2d ago A 27 tokens GPL-3.0

architect

05

cloudrift-ai/emmy

Agent Claude Code

Senior Architect reviewer. Reviews plans and diffs for simplicity, duplication, encapsulation and abstraction. Read-only — never edits code. Use before opening a PR, or when a design decision needs a second opinion.

80 yesterday A 44 tokens original Apache-2.0

discover-models

06

cloudrift-ai/emmy

Agent

Refresh the Emmy model lifecycle through bounded, read-only research.

80 yesterday A 11 tokens original Apache-2.0

onboard-model

07

cloudrift-ai/emmy

Agent

Qualify one model on the workflow-owned GPU and produce reviewed Emmy artifacts.

80 yesterday A 14 tokens original Apache-2.0