Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add skills/nick-pape/grackle/test-copilot-runtimenpx skills add nick-pape/grackle --skill test-copilot-runtimegit clone --depth 1 https://github.com/nick-pape/grackleWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/skills/nick-pape/grackle/test-copilot-runtime)<a href="https://agentmods.dev/skills/nick-pape/grackle/test-copilot-runtime"><img src="https://agentmods.dev/badge/skills/nick-pape/grackle/test-copilot-runtime.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00059 | $0.00709 |
| Opus 5 | $0.00030 | $0.00354 |
| Sonnet 5 | $0.00012 | $0.00142 |
| Haiku 4.5 | $0.00006 | $0.00071 |
Grade A, and why
test-copilot-runtime scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 4d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
How it starts
The opening of the file, as written. The whole thing — 47 lines — stays where its author put it; the contents beside it link to each section on GitHub.
Test the GitHub Copilot runtime
How to spawn the copilot runtime against an isolated test server. Assumes a server from /launch-grackle (with GRACKLE_URL + GRACKLE_API_KEY exported).
Allowed model names ✅ / ❌
| Model | Works? | Notes |
|---|---|---|
claude-sonnet-4.5 |
✅ confirmed live | The @github/copilot-sdk README default; spawns and runs |
gpt-5 |
⚠️ likely | Documented for Copilot CLI; not individually confirmed here |
claude-sonnet-4 |
⚠️ likely | From GitHub's supported-models docs |
claude-haiku-4.5 |
⚠️ likely | From GitHub's supported-models docs |
gpt-4o |
❌ does NOT work | Model "gpt-4o" is not available. GPT-4o is intentionally absent from the Copilot CLI/SDK (only in Copilot Chat + the VS Code extension) |
⚠️
grackle runtimesadvertisesgpt-4ofor copilot — that is wrong/unavailable for the CLI. Useclaude-sonnet-4.5. (Catalog entry:packages/common/src/runtime-catalog.ts; worth fixing — see the runtime-catalog.)
Spawn it
grackle persona create "Copilot Tester" --runtime copilot --model claude-sonnet-4.5 \
--prompt "You are a test agent. When asked to run a command, use your shell tool to run exactly that one command and nothing else."
grackle spawn local "Use your shell tool to run exactly this one command, then stop: cat /nonexistent_file_xyz" --persona copilot-tester
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 4d ago First seen · 47 lines · 0 tokens per session scan A 587b3e7d140b
test-copilot-runtime is a skill published in the GitHub repository nick-pape/grackle (21 stars, last pushed 2mo ago), licensed MIT. It adds 59 tokens to every session and 709 once invoked, about $0.0003 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other skills, from other repositories
copilot-pr-review-loop
Drive a GitHub pull request through repeated rounds of Copilot code review until convergence. Use when the user asks to "request Copilot review", "run a Copilot review loop", iterate on Copilot feedback, or wants automated triage-and-respond on Copilot PR comments. Covers re-request mechanics, open-thread filtering…
upstream-sync
Periodically sync new commits from microsoft/terminal into this manually-forked intelligent-terminal repo by cherry-picking commit-by-commit onto a dated sync branch, auto-skipping revert pairs and empty commits, auto-resolving known take-upstream files, and stopping cleanly on genuine conflicts. The agent (you…
release-notes
Generate user-facing release notes for Intelligent Terminal. Use when asked to write release notes, changelog, what-is-new summary, or prepare a release. Compares git commits between releases, looks up PR-linked issues and community contributors, then outputs formatted notes with "Verbed + Impact + Scenario" style…
add-acp-agent-support
Add first-class support for an ACP-compatible agent CLI to Intelligent Terminal. Use when integrating a new built-in AI agent, ACP server command, authentication flow, model selection, interactive delegation, session hooks, onboarding, Settings, branding, GPO policy, documentation, tests, build, deployment, or live…
opentag
Use when installing, pairing, operating, or troubleshooting OpenTag through the published CLI, self-hosted Control Plane, governed completion, supported collaboration platforms, or built-in coding agents.
pr-integration-test
Design, implement, and validate Intelligent Terminal integration tests for a target pull request or regression. Use when asked to add PR integration tests, convert a bug fix into E2E coverage, prove existing behavior still works, map tests to the release checklist, or verify E2E reports mark checklist cases complete.