cli-control

A testing guide for terminal programs, including command-line tools and text-based interfaces. It gathers repeatable local evidence about their output and interactive behavior without changing production code.

In plain words
What is it for?
Use it to verify normal command output and status, drive interactive prompts, test interrupts and cleanup, check terminal restoration, or investigate hangs and resizing behavior.
Why use it?
It helps reveal problems with prompts, input, signals, hangs, window resizing, startup, exit codes or terminal cleanup. It uses the smallest suitable test method and records only the evidence needed.

Skill for Claude CodeCodex

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/yarlson/yarstack/cli-control
Any agent
npx skills add yarlson/yarstack --skill cli-control
Clone the repo
git clone --depth 1 https://github.com/yarlson/yarstack

Made for: Claude Code, Codex.

Per session 42 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 295 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5 $0.00042 $0.00295
Opus 5 $0.00021 $0.00148
Sonnet 5 $0.00008 $0.00059
Haiku 4.5 $0.00004 $0.00030

Measured yesterday against content hash e117f36d8754, method: parsed. Prices are Anthropic first-party input rates as of 2026-08-30, from the pricing page.

Security

Grade A, and why

cli-control scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

plugins/yarstack/skills/cli-control/SKILL.md · 24 lines

What it actually says

CLI Control

Gather reproducible evidence from terminal-hosted interfaces without changing production behavior.

Workflow

  1. Identify the command, working directory, terminal conditions, and observable contract.
  2. Choose the smallest repeatable method: direct execution or repository tests for ordinary output and status; a terminal-capable harness for prompts, signals, resizing, hangs, or restoration.
  3. Prefer an existing trusted harness. Otherwise use a temporary local harness outside the repository without adding dependencies.
  4. Drive one action at a time and capture the smallest transcript or measurement that proves or disproves the behavior.
  5. Check interrupts, cleanup, exit status, hangs, and terminal restoration only where relevant.
  6. Clean up only sessions, processes, and artifacts created by this workflow.

Keep verification read-only unless the parent task separately authorizes implementation. Use behavior-implement and test-design for fixes or regression tests. Leave browser, desktop, Electron, and other graphical surfaces to ui-control.

Do not send credentials or destructive commands, retain sensitive transcripts, or keep one-off harness code in the repository without a product-test consumer.

Finish with the command and conditions, observed output and exit status, evidence, and any blocker or limitation.

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. yesterday First seen · 24 lines · 42 tokens per session scan A e117f36d8754

Subscribe to this mod's changes

cli-control is a skill published in the GitHub repository yarlson/yarstack (3 stars, last pushed 4d ago), licensed MIT. It adds 42 tokens to every session and 295 once invoked, about $0.0002 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-31.

Related

Other skills, from other repositories

opensrc

Fetch dependency source code to give AI agents deeper implementation context. Use when the agent needs to understand how a library works internally, read source code for a package, fetch implementation details for a dependency, or explore how an npm/PyPI/crates.io package is built. Triggers include "fetch source for"…

vercel-labs/opensrc · 103 tokens

release-cli

Release a new version of the tokenmaxxing CLI to npm. Use when explicitly asked to release or publish @851-labs/tokenmaxxing, including checking main, bumping apps/cli/package.json, updating CHANGELOG.md, committing, tagging cli-vX.Y.Z, pushing, monitoring generated native package publishing, and smoke testing the…

851-labs/tokenmaxxing · 73 tokens

ship-mobile-app

Build, change, debug, and prepare production mobile apps across domain meaning, local and server state, lifecycle, platform, and release boundaries. Use for non-trivial Flutter, React Native, iOS, or Android work involving persistence, sync, auth, offline behavior, notifications, permissions, native integrations…

aiopshwang/ship-mobile-app · 103 tokens

verify-regression-tests

Verify that a regression test actually detects the defect it claims to guard against. Use after adding or reviewing a bug-fix guard, reproducing the original failure after a fix, or investigating a suspiciously green regression test. Do not use for general TDD, broad test-suite audits, mutation-score optimization…

aiopshwang/verify-regression-tests · 75 tokens

repository-bootstrap

Audit and establish or upgrade the canonical repository governance baseline. Use only when explicitly asked to bootstrap, standardize, or upgrade a repository; do not use for feature development or source-code restructuring.

cz1993/vibe-coding-repository-standard · 41 tokens

repository-hygiene

Perform a read-only evidence-based audit of stale docs, agent files, scripts, generated artifacts, duplicate code, and optional tools. Use only when explicitly asked for cleanup or governance review; do not delete or refactor during the audit.

cz1993/vibe-coding-repository-standard · 52 tokens