Getting it into your agent
One page per mod, every tool's command on it. A separate URL per tool would split the same page into five that compete with each other.
npx agentmods add commands/benchbox-dev/benchbox/prgit clone --depth 1 https://github.com/BenchBox-dev/BenchBoxWrote this? Show the measurements
A badge with what this costs and how it scanned, read live from this page, so it follows the numbers instead of freezing them. Markdown for a README, HTML for a documentation site or a project page.
[](https://agentmods.dev/commands/benchbox-dev/benchbox/pr)<a href="https://agentmods.dev/commands/benchbox-dev/benchbox/pr"><img src="https://agentmods.dev/badge/commands/benchbox-dev/benchbox/pr.svg" alt="Measured on agentmods" height="20"></a>What it costs to keep this loaded
Counted locally with the o200k_base tokenizer, which is exact for GPT models; Claude uses its own tokenizer and its counts differ. Treat this as one consistent yardstick across the catalogue rather than a bill. Prices are per million input tokens.
| Model | Per session | Once invoked |
|---|---|---|
| Fable 5.1 | $0.00027 | $0.00417 |
| Opus 5 | $0.00014 | $0.00209 |
| Sonnet 5 | $0.00005 | $0.00083 |
| Haiku 4.5 | $0.00003 | $0.00042 |
Grade A, and why
pr scanned grade A with 0 findings against 26 rules in 11 categories — prompt injection, anti-refusal, data exfiltration, privilege escalation, supply chain, agent snooping, system-prompt leakage, SSRF and excessive agency — measured yesterday.
A static scan of the body, not an audit. Every finding is printed with the line that produced it so you can judge whether it matters here. A mod is markdown that instructs an agent; that is exactly why what it instructs is worth reading.
Nothing flagged
None of the 26 patterns this scan looks for appear in this file: no shell pipes, no recursive deletes, no credential paths, no hidden text, no instruction-override or anti-refusal phrasing, no agent-config snooping. That is not a guarantee, it is the absence of the things that are checkable.
What it actually says
Context
- Branch: !
git branch --show-current - Status: !
git status --short - Last commit: !
git log -1 --oneline - Ahead/behind develop: !
git rev-list --left-right --count origin/develop...HEAD 2>/dev/null || echo "(no origin/develop)"
Your task
BenchBox PR workflow targeting develop with linear history and squash-only merging.
Execute the following in order, stopping on the first failure:
-
Run
make agent-write-preflight. If it refuses (BenchBox primary clone), stop and tell the user to create a worktree (make worktree-create BRANCH=<name> WORKTREE_PATH=<path>). Do not commit, push, or open a PR from the primary clone withoutBENCHBOX_ALLOW_MAIN_CLONE_WRITE=1. -
Refuse if on
developormain. Stop and switch to a feature branch worktree if needed. -
Stage authorized changes and commit. Stage authorized paths explicitly (never
git add -A), verifymake agent-identity-check, and create a conventional commit. -
Run
make pr-preflightas the path-aware local gate. CI-only coverage remains separate. Fix root causes if failing; do not use--no-verify. -
Run
make pr-open— pushes branch and opens a PR vsdevelop. Note: auto-merge stays withheld unlessREADY=1(make pr-open READY=1ormake pr-ready). -
Print the PR URL and reported auto-merge status. Do not poll CI: pending is terminal.
What this file has done since we first saw it
Hashed on every crawl. A supply-chain change to an agent config is a question of when, not whether, so the history is kept rather than the latest state alone.
- yesterday First seen · 35 lines · 27 tokens per session scan A 157e0ca2b129
pr is a command published in the GitHub repository BenchBox-dev/BenchBox (12 stars, last pushed yesterday), licensed MIT. It adds 27 tokens to every session and 417 once invoked, about $0.0001 per session on Opus 5. A static security scan graded it A with 0 findings. No closer match exists in the catalogue, so it is treated as the original; first seen 2026-09-04.
Other commands, from other repositories
cr
When user uses /cr slash command, review uncommitted code changes with reusable language-aware rules, preferring project-local review commands or rules when they exist.
standup
Review my git commits from the past workday and generate a standup summary.
git
The pre-finish status: branch, hygiene findings, message checks, workflow lint, template state.
commit-claude-config
Version the Solana AI Kit config in git (un-ignores .claude/, CLAUDE.md, .mcp.json, .gitmodules and commits them).
commit
智能生成 Git 提交信息并提交.
session-report
Capture what changed this session and why, scoped to the current branch. Read by ship verbs when synthesizing the commit message; deleted after a successful commit.