pr
01Command Claude Code
BenchBox PR workflow - path-aware preflight, push, open PR vs develop; does not enable auto-merge unless READY=1.
Command Claude Code
BenchBox PR workflow - path-aware preflight, push, open PR vs develop; does not enable auto-merge unless READY=1.
Settings file Claude Code
Agent settings configuring enabledMcpjsonServers, extraKnownMarketplaces, enabledPlugins.
Skill Claude Code
Use when the user asks to "test TPC-H", "check compliance", "review architecture", "run quality checks", "check binaries", "test dialect translation", "compare implementations", "run live platform tests", "cut a release", "finalize a release", or "plan and execute" a benchmark feature.
Skill Claude Code
Organize and execute complex multi-step work through an executive, persistent named managers, focused workers, and independent review. Use when work divides across parallel workstreams or requires an independent review gate.
Skill Claude Code
Use for "implement code", "build a feature", "refactor code", "commit code", "review code", "adversarially review code", "review a code change", "review all code work in this session", "address PR review follow-ups", "run a PR review follow-up sweep", "clear PR backlog", "process PR backlog", "fix lint/type error"…
Skill Claude Code
Use when the user asks to "create documentation", "build docs", "review docs", "compare documents", "compress docs", "adversarial review docs", or "commit docs".
Skill Claude Code
Select model tiers, map reasoning effort, and dispatch delegated work through native or external agent harnesses. Use when a workflow must choose an agent model or effort level, or launch a manager, worker, or independent reviewer; do not use for direct, undelegated tool calls.
Skill Claude Code
Unified source-code selection and change-execution workflow: reuse ladder, vertical slicing, post-edit verification, named branches, commits, and authorized-write PRs.
Skill Claude Code
Unified investigation workflow: comparing artifacts, pre-edit research, context trust/authority handling, root-cause debugging, and validation-driven compression.
Skill Claude Code
Shared protocol for review-shaped actions, authorization scope, defect routing, solution-fit assessment, L1/L2/L3 planning-depth layers, local-only capture, and plan prior-decision reconciliation.
Skill Claude Code
Use when the user asks to sync, set up, inspect, validate, verify, pin, prune, promote, or configure skills managed by skill-sync.
Skill Claude Code
Use when the user asks to "run tests", "create tests", "fix failing test", "add test coverage", "fix slow tests", or "commit test changes".
Skill Claude CodeCodex
Consolidate accumulated permission grants across Claude Code, Codex, and Gemini: move trusted commands into project settings, clean garbage entries, verify cross-agent consistency, commit project-level configs.
MCP server Claude CodeCodexCursor +2
MCP server "todo-db" as configured in BenchBox-dev/BenchBox. Runs locally from the _project/scripts Python package. Needs 1 environment variable to run.
Instructions file Claude Code
Claude Code instructions for BenchBox-dev/BenchBox: Read and follow AGENTS.md; it is the active BenchBox authority. Load relevant generated skills from .claude/skills/.
Instructions file Gemini CLI
Gemini CLI instructions for BenchBox-dev/BenchBox: Read and follow AGENTS.md; it is the active BenchBox authority. Load relevant generated skills from .agents/skills/. Keep the shared mirror at that path; .gemini/skills/ is no longer populated.
Skill Claude Code
Use when working in a project tracked by todo-db — "what should I work on", "what's ready", "claim a TODO", "start work on an item", "record progress", "finish an item", "create a TODO", "add a work item", "check scope", "release my claim", "why did finish fail", "tracker stats", "prioritize TODOs", "batch…
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: