Use at the start of any non-trivial engineering task to load the right agent-framework doc(s) for the work — code review or PR prep, security review, pentest, hardening, threat modeling, supply-chain/dependency risk, debugging, quality gates, test-evidence ownership, harness routing/delegation, cost-first model…
Route substantive coding, planning, architecture, debugging, review, security, and research across the active coordinator, an independent external model family such as OpenCode/GLM, and optional fast, balanced, or deep Codex peers. Use when adaptive multi-model orchestration is adopted or requested, including deciding…
Orchestrate bounded, default-off work between cmux (local UI/session transport and lifecycle) and Hermes (provider router, plan/delegation brain, fallback, and usage ledger on a Tailscale-only VPS) using a local deterministic broker. Use when an operator has adopted cmux as the local surface and Hermes as the remote…
Arbitrate between competing plans or execution lanes on evidence, capability, budget, and risk instead of first-past-the-post. Use when two or more plans, agents, or model lanes could plausibly own a task and the coordinator must pick the efficient-frontier handoff, watchdog-verify delegated results, and enforce…
Gate and use optional repository context accelerators such as Graphify-compatible graphs/context maps, OpenWiki-compatible generated agent wikis, symbol indexes, or code-review graphs. Use when the user or repo opts into one of these tools, asks for token-saving context strategy, scope mapping, blast-radius analysis…
Route into the albrand/ai-config-kit operating framework and load only the relevant doctrine for planning, implementation, debugging, review, security, harness design, quality convergence, repository adoption, journaling, skill learning, token economy, or context acceleration. Use when the user asks to apply, inspect…
Perform a whole-project assessment, generate or refresh knowledge docs, derive missing requirements and tickets from internal and external sources, and plan hardening work. Use when the user wants an existing project read deeply and turned into docs, gaps, QA logic, and implementation plans.
Turn business requirements, docs, designs, existing boards, legacy roadmap material, and stakeholder prompts into an evidence-backed roadmap, operating model, and PR-sized ticket ecosystem. Use when the user wants AI to bootstrap, import, reconcile, or create product/project roadmap artifacts for a new or legacy app.
Bootstrap or reconcile the technology side of a project from a stack prompt, existing repos, infrastructure choices, cloud/provider access, local environment needs, CI/CD gates, PR automation, and AI developer workstreams. Use when the user wants a project-agnostic technical framework setup plan or live setup actions.
Turn a product or platform module idea into an evidence-backed delivery plan using AI analysis plus normal repository, documentation, and board inspection. Use when a user provides a module name, capability area, roadmap item, Linear/project-planning request, or migration target and wants Codex or Claude Code to study…
Discover and use the host software an agent runs in or through — local session transports (cmux), terminal multiplexers (tmux, zellij), and other agentic shells/harnesses — via a capability-first, host-neutral adapter contract. Use when an agent needs to detect which native surfaces are present, route on declared…
Multi-pass, multi-lens security review that reinforces detection through independent blind finder passes and an adversarial refute pass. Use when reviewing an owned or authorized codebase/change for vulnerabilities with high confidence, or when a change touches auth, data, crypto, external input, dependencies, or…
Authorization-gated white-hat penetration testing against an owned or explicitly authorized target. Follows find → validate → fix → regress. Requires the authorization gate before any active testing (crafted requests, exploit runs, scanning, fuzzing). Use for "pentest my app", "test this endpoint for vulnerabilities"…
Find and load the right Codex skill from a large local library when skill descriptions are shortened, hidden, explicit-only, or no visible skill clearly matches. Use before giving up on skill context, after skill or plugin library changes, or when Codex needs smart access to all installed skills without injecting…
UX/product workflow for design-makers and consumers of existing designs. Use for making or revising mockups, screens, layouts, flows, prototypes, tokens, typography, components, design systems, live design previews, signoff, and Figma or board handoff. Also use to consume a prototype, Figma file, screenshots, or UI…
Use when writing or debugging a bb plugin — imports failing at load, a migration corrupting state, an RPC call rejected before it leaves the browser, or a build that works in a shell and not in bb.
Use when calling the bb plugin SDK — listing providers or threads, reading thread events, spawning threads on a machine, or creating worktrees. Several calls return confident, wrong-looking-right answers; this is what each one actually does.
Use when delegating execution to GLM, opencode, or any non-Anthropic agent — choosing how to invoke it, or wondering why a provider you definitely used shows zero usage in a cost report. There are two ways to call it and only one of them is accountable.
Use when a coding agent on a headless Linux box or VPS FAILS — every shell command denied, "setting up uid map: Permission denied", a sandbox or bwrap error, a login that hangs with no code printed, EACCES on a global npm install, or a daemon that dies when SSH disconnects. Host policy masquerading as agent bugs.
Use when building or debugging token/cost accounting across agent providers, or when a spend dimension renders empty. Covers what each event source can and cannot tell you, why skill usage is nearly invisible, and how to bucket a usage chart so it is readable at every window.
Use when running a fleet of agents, or when the same class of mistake keeps recurring across sessions. Failure patterns that repeat, why each one survives review, and the check that catches it. Read before declaring work done.
★not rated 5 yesterdayA50 tokens
At most 3 mods per repository are shown here, and a mod shipped inside a plugin is left to that plugin's page — the rest are on their repository pages: