delegating-to-local-llm

delegating-to-local-llm is a skill for Claude Code from isvlasov/rageatc-oss. It costs 102 tokens per session (990 once invoked), scanned A, original, MIT.

A procedure for handing coding work to a local language model running on your own Apple-silicon Mac. The model works in a visible session that you can monitor and steer.

In plain words
What is it for?
Use it to check local-model availability, choose a downloaded model, start a coding subagent, send it tasks, and monitor or redirect its session.
Why use it?
It can move suitable tasks away from a cloud model and keep the work running offline, without using cloud-model tokens. You still need to review the result.

Skill for Claude Code

Written for Claude Code: shipped in a Claude Code plugin. Also seen: mentions subagents; mentions Claude Code.

Part of the rageatc-code-oss plugin — 21 skills, 5 agents, 1 hook shipped together

Good fit Use it to check local-model availability, choose a downloaded model, start a coding subagent, send it tasks, and monitor or redirect its session.

Compare 6 skills from other repositories ↓
Install with agentmods
npx agentmods add skills/isvlasov/rageatc-oss/delegating-to-local-llm
Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

Any agent
npx skills add isvlasov/rageatc-oss --skill delegating-to-local-llm
Clone the repo
git clone --depth 1 https://github.com/isvlasov/rageatc-oss

Made for: Claude Code.

Or install rageatc-code-oss, the plugin that ships this one along with the rest of its 21 skills, 5 agents, 1 hook.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for delegating-to-local-llm

README.md
[![agentmods](https://agentmods.dev/badge/skills/isvlasov/rageatc-oss/delegating-to-local-llm/github.svg)](https://agentmods.dev/skills/isvlasov/rageatc-oss/delegating-to-local-llm)
Your own site
<a href="https://agentmods.dev/skills/isvlasov/rageatc-oss/delegating-to-local-llm"><img src="https://agentmods.dev/badge/skills/isvlasov/rageatc-oss/delegating-to-local-llm/github.svg" alt="Measured on agentmods" height="20"></a>

Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.

agentmods 80×15 button for delegating-to-local-llm

Your own site · 80×15
<a href="https://agentmods.dev/skills/isvlasov/rageatc-oss/delegating-to-local-llm"><img src="https://agentmods.dev/badge/skills/isvlasov/rageatc-oss/delegating-to-local-llm.svg" alt="Reviewed on agentmods" width="80" height="20"></a>
Per session 102 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 990 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 1 finding. A grade says what 26 rules found in the file — not that it is safe.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00102 $0.00990
Opus 5 $0.00051 $0.00495
Sonnet 5 $0.00020 $0.00198
Haiku 4.5 $0.00010 $0.00099

Measured 12d ago against content hash f260c2f98bff, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-11, from the pricing page.

Security

Grade A, and why

delegating-to-local-llm scanned grade A with 1 finding against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 12d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Makes network callslowCapability

Not a fault in itself. Listed so you know the mod talks to something, and to what.

curl -s --max-time 3 -H "Authorization: Bearer $KEY" http://127.0.0.1:8000/v1/models
plugins/rageatc-code-oss/skills/delegating-to-local-llm/SKILL.md · 53 lines

How it starts

The opening of the file, as written. The whole thing — 53 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Delegating to a Local LLM

The subagent is an interactive Pi session pinned to a local omlx model, spawned in a herdr pane. herdr provides the whole control loop — spawn, state, read, send — and Pi's herdr integration reports working/blocked/idle automatically. The session stays visible and steerable; you own acceptance of its work.

Preflight

uname -m                                                  # arm64 required (omlx is Apple-Silicon MLX)
KEY=$(jq -r .auth.api_key ~/.omlx/settings.json)
curl -s --max-time 3 -H "Authorization: Bearer $KEY" http://127.0.0.1:8000/v1/models
which pi && herdr agent list >/dev/null && echo ok
  • curl fails -> omlx start; succeeds but lists no models -> models need downloading
  • pi or herdr missing, not arm64, or first run on this machine -> walk the user through references/setup.md

Delegate

  1. Pick the model from the /v1/models response. The user's named choice wins; otherwise ask which model they prefer for the task. Memory headroom is model- and machine-specific: if omlx's prefill guard aborts mid-task ("Prefill context too large for available memory"), the model cannot handle the accumulated context on this machine — restart the task on a lighter model.

  2. Spawn — short kebab task name; cwd is the project the task touches. Pin the pane to your own workspace — spawning defaults to whichever workspace the user has focused at that moment, which may not be yours:

    WS=$(herdr pane get "$HERDR_PANE_ID" | jq -r '.result.pane.workspace_id')
    herdr agent start <task-name> --cwd <dir> --workspace $WS --no-focus -- omlx launch pi --model <model-id> --api-key $KEY
    

    Capture pane_id from the JSON response. (omlx launch also rewrites Pi's omlx provider config, so it is always current.)

  3. Wait ready: herdr agent wait <task-name> --status idle --timeout 30000

  4. Send the task, then submit it:

    herdr agent send <task-name> "<task>"
    sleep 1 && herdr pane send-keys <pane-id> enter
    

Read the full file on GitHub · 53 lines

Files

What ships with it

1 file beside SKILL.md in the same directory: the scripts, references and assets a skill reads on demand. Not counted in the per-session cost; read them before you install if any of them is executable.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 12d ago First seen · 53 lines · 102 tokens per session scan A f260c2f98bff

Subscribe to this mod's changes

delegating-to-local-llm is a skill published in the GitHub repository isvlasov/rageatc-oss (9 stars, last pushed 1mo ago), licensed MIT. It adds 102 tokens to every session and 990 once invoked, about $0.0005 per session on Opus 5. A static security scan graded it A with 1 finding (makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.