Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx skills add airiclenz/llama-launcher --skill manage-llm-servergit clone --depth 1 https://github.com/airiclenz/llama-launcherWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/airiclenz/llama-launcher/manage-llm-server)<a href="https://agentmods.dev/skills/airiclenz/llama-launcher/manage-llm-server"><img src="https://agentmods.dev/badge/skills/airiclenz/llama-launcher/manage-llm-server/github.svg" alt="Measured on agentmods" height="20"></a>Or the 80×15 button, for a site that already has a row of RSS and ATOM ones. Only the verdict fits; the numbers stay here.
<a href="https://agentmods.dev/skills/airiclenz/llama-launcher/manage-llm-server"><img src="https://agentmods.dev/badge/skills/airiclenz/llama-launcher/manage-llm-server.svg" alt="Reviewed on agentmods" width="80" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00066 | $0.01738 |
| Opus 5 | $0.00033 | $0.00869 |
| Sonnet 5 | $0.00013 | $0.00348 |
| Haiku 4.5 | $0.00007 | $0.00174 |
Grade A, and why
manage-llm-server scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 11d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 96 lines — stays where its author put it; the contents beside it link to each section on GitHub.
manage-llm-server
llama-launcher is a CLI for running and switching local LLM servers (llama.cpp, Ollama, LM Studio). Models are defined as profiles; loading a profile starts the right server and loads its model. Use this skill to inspect state, switch profiles, and tail logs.
Safety rules (do these first)
- Always start with a passive inspection —
llama-launcher status --jsonandllama-launcher list --json. Do not callload,unload,start, orstopuntil you've shown the user what's currently running and have explicit confirmation to change it. - Check for in-flight work before mutating. Switching, unloading, or stopping a model interrupts any request or job currently using it. If a high
uptime_secondscoincides with the user actively working, assume the model may be in use and confirm before changing it. load <profile>is idempotent — it's a no-op if the requested profile is already active. Only reach for--restartif the user explicitly asks to force a restart, or if the server is in a bad state confirmed via logs.unloadandstopend the session — any in-flight requests will fail. Confirm before calling.
Command reference
Passive (safe to run any time):
| Command | Purpose |
|---|---|
llama-launcher list / --json |
Enumerate available profiles (name, title, description, backend, model file, gpu_layers, context_size). title is the human-readable label; description is omitted when unset. |
llama-launcher status / --json |
Per-backend: running, address, active_profile, active_model, pid, uptime_seconds. |
llama-launcher logs <backend> |
Tail the last chunk of an instance's log. Add -f to follow. |
llama-launcher logs clean --days N / --all |
Prune old logs. |
llama-launcher config validate |
Check the config file for errors. |
llama-launcher version |
Print version. |
Mutating (confirm with user first):
| Command | Purpose |
|---|---|
llama-launcher load <profile> |
Activate a profile. No-op if already active. |
llama-launcher load <profile> --restart |
Force a restart even if active. |
llama-launcher unload [profile] |
Stop the server / unload the model. |
llama-launcher start [--profile p] |
Start a server, optionally with a profile. |
llama-launcher stop [target] |
Stop a server by host:port or backend name. |
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 11d ago First seen · 96 lines · 66 tokens per session scan A 009c45535ed9
manage-llm-server is a skill published in the GitHub repository airiclenz/llama-launcher (2 stars, last pushed 28d ago), licensed MIT. It adds 66 tokens to every session and 1,738 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.
Other skills, from other repositories
gog-workspace
Use the gog CLI for Google Workspace tasks across Gmail, Calendar, Drive, Docs, Sheets, Contacts, and related services. Use when the user asks to check email, search Gmail, inspect calendar events, find Drive files, read Docs or Sheets, or manage Google Workspace data through gog.
apple-calendar
Read macOS Calendar events via the icalBuddy CLI and create events via AppleScript (osascript). Use to check the user's calendar, agenda, upcoming events, or add an event on macOS.
apple-reminders
Manage Apple Reminders via the remindctl CLI on macOS — list, add, complete, delete, manage lists.
github
Drive GitHub via the official gh CLI — repos, issues, pull requests, releases, gists, Actions runs, and raw REST through gh api. Use when the user asks to inspect or manage GitHub.
obsidian
Read, search, and create Markdown notes inside an Obsidian vault on disk.
Manipulate PDF files — merge, split, extract pages/text, PDF↔images, OCR, info — via the qpdf / poppler / ocrmypdf CLIs. Use to combine, slice, convert, or OCR PDFs.