dmmdea/offload-harness

Delegate summarize/classify/extract/triage to a FREE local Gemma-4 cascade via llama.cpp. Go CLI + MCP server; never calls a cloud model. Includes a turnkey setup skill for Claude Code.

2Stars on the repository
5Mods indexed here, across every type
2d agoLast push, which is what freshness is scored on
Apache-2.0Licence, which decides whether bodies are shown

local-offload-setup

01

dmmdea/offload-harness

Skill Claude CodeCodex

Use when setting up the "local-offload" harness on a Windows machine — a free local Gemma-4 cascade that lets a coding agent delegate short-context grunt work (summarize / classify / extract / triage) so those tokens never hit the cloud context. Cross-vendor: NVIDIA (CUDA, ≥8GB), AMD Radeon incl. RDNA3 iGPUs like the…

2 2d ago A 204 tokens original Apache-2.0

pp-comfyui

02

dmmdea/offload-harness

Skill Claude CodeCodex

Drive a local ComfyUI render server from the shell, with a durable record of every run the server itself forgets. Trigger phrases: submit a comfyui graph, why can't comfyui see my model, how long did that render take, what made this output file, check what this loader accepts, use comfyui, run comfyui.

2 2d ago A 81 tokens original Apache-2.0

pp-llamaswap

03

dmmdea/offload-harness

Skill Claude CodeCodex

The llama-swap operations console — durable history, drain-aware control, and measurement commands with three specific guarantees: keep-set unloads are refused statically by id AND alias (never from server ttl), --drain fails closed when slot state is unreadable, and fit/ctx refuse to answer inside their uncertainty…

2 2d ago A 138 tokens original Apache-2.0