Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/rivet-dev/sandbox-agent/post-release-testinggit clone --depth 1 https://github.com/rivet-dev/sandbox-agentWhat it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5 | $0.00000 | $0.00339 |
| Opus 5 | $0.00000 | $0.00169 |
| Sonnet 5 | $0.00000 | $0.00068 |
| Haiku 4.5 | $0.00000 | $0.00034 |
Grade C, and why
post-release-testing scanned grade C with 2 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured 2d ago.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Downloads and executes remote codehighSupply chain
curl | sh runs whatever the server returns today, which is not necessarily what it returned when this was reviewed.
curl -fsSL https://releases.rivet.dev/sandbox-agent/0.4.x/install.sh | sh && Makes network callslowCapability
Not a fault in itself. Listed so you know the mod talks to something, and to what.
- Installs sandbox-agent via curl from releases.rivet.dev What it actually says
Post-Release Testing Agent
You are a post-release testing agent. Your job is to verify that a sandbox-agent release works correctly.
Environment Setup
First, source the environment file:
source ~/misc/env.txt
Tests to Run
Run these tests in order, reporting results as you go:
1. Docker Example Test
RUN_DOCKER_EXAMPLES=1 pnpm --filter @sandbox-agent/example-docker test
This test:
- Creates an Alpine container
- Installs sandbox-agent via curl from releases.rivet.dev
- Verifies the
/v1/healthendpoint responds correctly
2. E2B Example Test
pnpm --filter @sandbox-agent/example-e2b test
This test:
- Creates an E2B sandbox with internet access
- Installs sandbox-agent via curl
- Verifies the
/v1/healthendpoint responds correctly
3. Install Script Test
Manually verify the install script works in a fresh environment:
docker run --rm alpine:latest sh -c "
apk add --no-cache curl ca-certificates libstdc++ libgcc bash &&
curl -fsSL https://releases.rivet.dev/sandbox-agent/0.4.x/install.sh | sh &&
sandbox-agent --version
"
Instructions
- Run each test sequentially
- Report the outcome of each test (pass/fail)
- If a test fails, capture and report the error output
- Provide a summary at the end with overall pass/fail status
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- 2d ago First seen · 57 lines · 0 tokens per session scan C 301435d9f6ea
post-release-testing is a command published in the GitHub repository rivet-dev/sandbox-agent (1,551 stars, last pushed 2mo ago), licensed Apache-2.0. It costs nothing until one of its globs matches a file; then it loads 339 tokens. A static security scan graded it C with 2 findings (downloads and executes remote code, makes network calls). No closer match exists in the catalogue, so it is treated as the original; first seen 2026-08-30.
Other commands, from other repositories
MIGRATE_DESIGN
Design doc for the migration tool PR. Author: Sol ([email protected]). Co-authored-by: wakesync.
awesome-chatgpt
Search awesome-ChatGPT-repositories for open-source GitHub repositories related to ChatGPT and LLMs.
agentlas-cloud
Staff a task only from the signed-in owner's Agent Cloud agents.
commit
智能生成 Git 提交信息并提交.
tasks
Command "tasks" from thrashr888/agentkernel, covering durable tasks, use up to four task workers (the default) and bound active tasks explicitly.
capture-feedback
Quick feedback capture with structured signals.