test-acp-runtime

test-acp-runtime is a skill for Claude Code from nick-pape/grackle. It costs 64 tokens per session (784 once invoked), scanned A, original, MIT.

A test workflow for starting and exercising ACP agent runtimes against an isolated Grackle test server. ACP is a standard protocol for connecting coding agents to applications, and these runtimes include Claude Code, Codex, and Copilot variants.

In plain words
What is it for?
Use it to test Claude Code ACP, Codex ACP, or Copilot ACP against a test server and confirm their tool calls and model configuration work.
Why use it?
It helps verify how each runtime behaves and exposes real tool failures that some native runtimes may hide. It also documents the model settings needed to start each runtime.

Skill for Claude Code

Written for Claude Code: installed under .claude/. Also seen: reads .claude/ paths; mentions Codex.

Install

Getting it into your agent

One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.

agentmods
npx agentmods add skills/nick-pape/grackle/test-acp-runtime
Any agent
npx skills add nick-pape/grackle --skill test-acp-runtime
Clone the repo
git clone --depth 1 https://github.com/nick-pape/grackle

Made for: Claude Code.

Wrote this? Show the measurements

A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.

agentmods badge for test-acp-runtime

README.md
[![agentmods](https://agentmods.dev/badge/skills/nick-pape/grackle/test-acp-runtime.svg)](https://agentmods.dev/skills/nick-pape/grackle/test-acp-runtime)
Your own site
<a href="https://agentmods.dev/skills/nick-pape/grackle/test-acp-runtime"><img src="https://agentmods.dev/badge/skills/nick-pape/grackle/test-acp-runtime.svg" alt="Measured on agentmods" height="20"></a>
Per session 64 Skills are progressive disclosure: only the name and description are preloaded; the body loads when the skill is used.
When invoked 784 The whole file, excluding the scripts and references it only reads on demand.
Security scan A 0 findings. Scan, not verified.
Origin original No closer match found in the catalogue.
Token cost

What it costs to keep this loaded

Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.

ModelPer sessionOnce invoked
Fable 5.1 $0.00064 $0.00784
Opus 5 $0.00032 $0.00392
Sonnet 5 $0.00013 $0.00157
Haiku 4.5 $0.00006 $0.00078

Measured 6d ago against content hash 737b7fb8cf28, method: parsed. Prices are Anthropic first-party input rates as of 2026-09-06, from the pricing page.

Security

Grade A, and why

test-acp-runtime scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 6d ago.

A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.

Nothing flagged

None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.

.claude/skills/test-acp-runtime/SKILL.md · 46 lines

How it starts

The opening of the file, as written. The whole thing — 46 lines — stays where its author put it; the contents beside it link to each section on GitHub.

Test the ACP runtimes

How to spawn the ACP-variant runtimes — claude-code-acp, codex-acp, copilot-acp — against an isolated test server. Assumes a server from /launch-grackle (with GRACKLE_URL + GRACKLE_API_KEY exported).

Why ACP matters: the native runtimes (claude-code/copilot/codex) are slated to be deprecated in favor of these ACP variants. ACP is a standard protocol with a uniform tool_call_update status (completed/failed), so one adapter (runtime-acp) covers all agents — and unlike native claude-code (which emits synthetic empty tool_results and can't surface failures), claude-code-acp reports real tool failures.

Model names

ACP runtimes are listed as model "(agent-selected)" by grackle runtimes, but a persona still requires a model (spawn errors with FailedPrecondition: persona has no model configured otherwise):

Runtime Model to pass
claude-code-acp sonnet / opus / haiku (runs Claude underneath) — validated with sonnet
codex-acp a Codex model (same ChatGPT-account gating as native — use gpt-5.5; see /test-codex-runtime)
copilot-acp a Copilot model, e.g. claude-sonnet-4.5 (see /test-copilot-runtime)

Spawn it

grackle persona create "ACP Tester" --runtime claude-code-acp --model sonnet --prompt "You are a test agent."
grackle spawn local "Run exactly this one shell command and nothing else, then stop: cat /nonexistent_file_xyz" --persona acp-tester

What it proves (confirmed live 2026-05-29)

  • Spawns cleanly. The earlier Invalid permissions.defaultMode: auto failure (#1366) is fixed by #1370 — claude-code-acp now isolates its CLAUDE_CONFIG_DIR from your personal ~/.claude (which may carry the interactive-only defaultMode: "auto").
  • Surfaces real tool failures. A failing command (cat /nonexistent) produced a tool_result with content {is_ok:false} + "tool_error":true and the real Exit code 1 / No such file text — via the adapter's status === "failed"toolError mapping (packages/runtime-acp/src/acp.ts). This is the case native claude-code cannot express.

Read the full file on GitHub · 46 lines

Changes

What this file has done since we first saw it

Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.

  1. 6d ago First seen · 46 lines · 64 tokens per session scan A 737b7fb8cf28

Subscribe to this mod's changes

test-acp-runtime is a skill published in the GitHub repository nick-pape/grackle (21 stars, last pushed 2mo ago), licensed MIT. It adds 64 tokens to every session and 784 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.

Related

Other skills, from other repositories

copilot-pr-review-loop

Drive a GitHub pull request through repeated rounds of Copilot code review until convergence. Use when the user asks to "request Copilot review", "run a Copilot review loop", iterate on Copilot feedback, or wants automated triage-and-respond on Copilot PR comments. Covers re-request mechanics, open-thread filtering…

microsoft/intelligent-terminal · 91 tokens

upstream-sync

Periodically sync new commits from microsoft/terminal into this manually-forked intelligent-terminal repo by cherry-picking commit-by-commit onto a dated sync branch, auto-skipping revert pairs and empty commits, auto-resolving known take-upstream files, and stopping cleanly on genuine conflicts. The agent (you…

microsoft/intelligent-terminal · 139 tokens

release-notes

Generate user-facing release notes for Intelligent Terminal. Use when asked to write release notes, changelog, what-is-new summary, or prepare a release. Compares git commits between releases, looks up PR-linked issues and community contributors, then outputs formatted notes with "Verbed + Impact + Scenario" style…

microsoft/intelligent-terminal · 73 tokens

add-acp-agent-support

Add first-class support for an ACP-compatible agent CLI to Intelligent Terminal. Use when integrating a new built-in AI agent, ACP server command, authentication flow, model selection, interactive delegation, session hooks, onboarding, Settings, branding, GPO policy, documentation, tests, build, deployment, or live…

microsoft/intelligent-terminal · 72 tokens

pr-integration-test

Design, implement, and validate Intelligent Terminal integration tests for a target pull request or regression. Use when asked to add PR integration tests, convert a bug fix into E2E coverage, prove existing behavior still works, map tests to the release checklist, or verify E2E reports mark checklist cases complete.

microsoft/intelligent-terminal · 66 tokens

opentag

Deploy or operate OpenTag's supported paired setup when a user needs to bootstrap the self-hosted Docker Compose Control Plane and Slack Source App, configure or pair a local ACP Runner with a GitHub Project Target, start the Runner service, verify readiness, or diagnose a Slack mention that did not complete.

amplifthq/opentag · 64 tokens